CtrlK
BlogDocsLog inGet started
Tessl Logo

statistical-experimental-evaluation

Design and run statistical experiments that test the formal problem, proposed methods, theoretical predictions, baselines, and ablations.

57

Quality

66%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./external/agents/stat_research_agent/skills/statistical-experimental-evaluation/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

72%Weight 40%Scale 1-3

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is token-efficient and well-organized with useful schema templates, but it stops at specification rather than giving executable steps or validation feedback loops for the experimental workflow.

Suggestions

Add an explicit sequenced workflow (e.g., 1. write config, 2. run methods, 3. collect metrics, 4. validate against claims) with validation checkpoints before recording verdicts.

Provide at least one executable snippet or command for running a condition and emitting metrics.json so the guidance is copy-paste ready.

DimensionReasoningScore

Conciseness

The body is lean and efficient: terse bullets, artifact paths, and compact JSON schemas with no concept-explaining fluff, assuming Claude's competence throughout.

3 / 3

Actionability

Provides concrete templates (metrics.json and claim_verdicts schemas, artifact paths) but no executable commands or code to actually run experiments; guidance is a spec rather than runnable instruction.

2 / 3

Workflow Clarity

Sections list what to define and produce (plan, artifacts, schema, rules) but there is no explicit execution sequence and no validation checkpoints for a batch/risky experimental process.

2 / 3

Progressive Disclosure

A single, well-organized SKILL.md with clear sections and no nested external references; no bundle files are needed or referenced, so the structure is appropriate.

3 / 3

Total

10

/

12

Passed

Description

60%Weight 40%Scale 1-3

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific about the experimentation domain and its targets, but lacks an explicit trigger/use-when clause and leans on somewhat technical jargon. Adding natural user-facing triggers would raise completeness and trigger-term quality.

Suggestions

Add an explicit 'Use when...' clause naming natural triggers (e.g., 'Use when running experiments, simulations, or comparing methods against baselines and ablations').

Include more colloquial trigger terms users would actually say rather than only academic phrasing like 'theoretical predictions' and 'formal problem'.

DimensionReasoningScore

Specificity

Lists multiple concrete actions ('Design and run statistical experiments that test...') targeting specific artifacts (formal problem, proposed methods, theoretical predictions, baselines, ablations).

3 / 3

Completeness

Clearly states what the skill does but lacks any explicit 'Use when...' trigger guidance, so completeness is capped per the rubric.

2 / 3

Trigger Term Quality

Contains relevant domain terms ('statistical experiments', 'baselines', 'ablations') but leans technical/academic; missing common natural variations a user might say.

2 / 3

Distinctiveness Conflict Risk

The statistical-experiment framing is a distinct niche, but the broad evaluation vocabulary could still overlap with general analysis or evaluation skills.

2 / 3

Total

9

/

12

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

Total

15

/

16

Passed

Repository
aiming-lab/AutoResearchClaw
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.