CtrlK
BlogDocsLog inGet started
Tessl Logo

analyze-results

Analyze ML experiment results, compute statistics, generate comparison tables and insights. Use when user says "analyze results", "compare", or needs to interpret experimental data.

70

Quality

87%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, well-structured instruction skill: lean, readable, and logically sequenced with no filler. Its only real gaps are the absence of a concrete output example (e.g., a sample comparison table) and explicit validation before results are reported.

Suggestions

Add a small example comparison table (or one concrete command snippet for reading JSON/CSV results) under Step 2 so the expected output shape is unambiguous.

Insert an explicit verification checkpoint after Step 2 (e.g., re-check that every run in the raw table made it into the comparison and that the baseline is correctly identified) before proceeding to statistics.

In Step 3, specify how to flag outliers (e.g., values beyond N standard deviations from the seed mean) so the instruction is executable without judgment calls.

DimensionReasoningScore

Conciseness

The body is lean and efficient — short bullets, no explanation of concepts Claude already knows, and every line instructs rather than describes, matching the 'every token earns its place' anchor.

5 / 5

Actionability

As an instruction-only skill its guidance is concrete ("Check figures/, results/", "Delta vs baseline", "mean +/- std", the Observation/Interpretation/Implication/Next step structure), but it stops short of fully executable specifics like an example table format or a concrete command, leaving minor gaps.

4 / 5

Workflow Clarity

Five steps are clearly sequenced with implicit checkpoints ("flag outliers", "check reproducibility"), but there is no explicit validation of the assembled comparison table or verification step before reporting findings, so it sits at 'most checkpoints present, minor validation gaps'.

4 / 5

Progressive Disclosure

The skill is under 50 lines, single-purpose, needs no external reference files (none exist in the bundle), and is organized into clearly labeled sections — the rubric's simple-skill exception applies.

5 / 5

Total

18

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it names the domain and several concrete capabilities and includes an explicit, natural-sounding 'Use when' clause. Its only weaknesses are slightly generic trigger coverage and the unqualified "compare" trigger.

DimensionReasoningScore

Specificity

"Analyze ML experiment results, compute statistics, generate comparison tables and insights" lists four distinct concrete actions with comprehensive coverage of the domain, matching the top anchor rather than the 'minor gaps' of a 4.

5 / 5

Completeness

It explicitly answers both 'what' (analyze results, compute statistics, generate tables and insights) and 'when' ("Use when user says 'analyze results', 'compare', or needs to interpret experimental data") with concrete trigger phrases.

5 / 5

Trigger Term Quality

Natural phrases like "analyze results", "compare", and "interpret experimental data" are present, but common synonyms (e.g., "compare models", "check results") and file extensions (.json/.csv) are missing, and bare "compare" is generic.

4 / 5

Distinctiveness Conflict Risk

"ML experiment results" carves a clear niche, but the bare trigger "compare" is generic enough to fire for unrelated comparison tasks, giving minor overlap risk with closely related skills.

4 / 5

Total

18

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.