CtrlK
BlogDocsLog inGet started
Tessl Logo

statistical-experimental-evaluation

Design and run statistical experiments that test the formal problem, proposed methods, theoretical predictions, baselines, and ablations.

63

Quality

75%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./external/agents/stat_research_agent/skills/statistical-experimental-evaluation/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

80%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A concise, well-structured skill body that gives concrete artifacts and schemas without padding. Its main weakness is the lack of explicit validation/checkpoint steps in the experimental workflow, which is important for batch statistical operations.

Suggestions

Add an explicit validation step in the workflow (e.g., 'Verify every claim_id in metrics.json maps to a formulated claim before interpreting results').

Include a short feedback loop: on failed runs, count them, diagnose, and re-run rather than silently dropping them.

Show one concrete runnable example (seed/folds/repetitions) so the artifact layout is immediately executable.

DimensionReasoningScore

Conciseness

Lean and efficient — assumes Claude's competence with no padding explaining what experiments or statistics are; every section (plan, artifacts, schema, rules) earns its place.

5 / 5

Actionability

Concrete guidance via an explicit artifact file layout and executable JSON schemas for metrics and claim verdicts, but it stops short of showing how to actually run an experiment, leaving a minor gap between structure and execution.

4 / 5

Workflow Clarity

A rough sequence exists (plan → artifacts → schema → rules) but there are no explicit validation checkpoints or feedback loops for batch/statistical operations, which caps workflow clarity at 3 per the rubric.

3 / 5

Progressive Disclosure

Well-organized into clearly labeled sections with no bundle files present and no need for external references; the content is self-contained and easy to navigate.

5 / 5

Total

17

/

20

Passed

Description

71%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A specific, action-oriented description with strong capability coverage, but it lacks an explicit 'when to use' clause, which limits its completeness and trigger usability. Adding a 'Use when...' sentence with natural keywords would raise it significantly.

Suggestions

Add an explicit trigger clause such as 'Use when designing experiments, running simulations, comparing methods against baselines, or producing statistical evidence'.

Include common user-facing synonyms in the description (simulation, evaluation, metrics, diagnostics) to improve trigger term coverage.

Mirror the trigger-keywords from metadata into the description so Claude can match natural user phrasing.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'Design and run statistical experiments that test the formal problem, proposed methods, theoretical predictions, baselines, and ablations' — covering the full experimental scope comprehensively.

5 / 5

Completeness

The 'what' is clearly stated but there is no 'Use when...' clause or equivalent explicit trigger guidance, which caps completeness at 3 per the judging guidelines.

3 / 5

Trigger Term Quality

Natural terms like 'statistical experiments', 'baselines', and 'ablations' are present, but common synonyms a user might say (e.g. 'simulation', 'evaluation', 'metrics') are absent from the description itself.

4 / 5

Distinctiveness Conflict Risk

The research-methodology niche (testing formal problems, theoretical predictions, baselines, ablations) is fairly distinct with low conflict risk, with only minor overlap with general evaluation skills.

4 / 5

Total

16

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

Total

15

/

16

Passed

Repository
aiming-lab/AutoResearchClaw
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.