CtrlK
BlogDocsLog inGet started
Tessl Logo

analyze-results

Analyze ML experiment results, compute statistics, generate comparison tables and insights. Use when user says "analyze results", "compare", or needs to interpret experimental data.

64

Quality

77%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/skills-codex/analyze-results/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A lean, well-structured workflow that assumes Claude's competence and never wastes tokens. Its weakness is actionability: key steps (parsing, delta computation, output table shape) are described at the level of intent rather than executable detail.

Suggestions

Add a short executable snippet or concrete command for parsing JSON/CSV result files (e.g., a pandas/json example), since Step 1 currently only names the directories to check.

Make "Delta vs baseline" executable by specifying the formula (e.g., (metric - baseline) / baseline) and how the baseline run is identified.

Show a minimal example of the required output format — a 2-3 row markdown comparison table plus one numbered finding — so the "Raw data table" and "Key findings" requirements are unambiguous.

DimensionReasoningScore

Conciseness

The ~40-line body is lean and efficient with no padding and no explanations of concepts Claude already knows; the only parentheticals ("perplexity, accuracy, loss") earn their place by disambiguating the domain.

5 / 5

Actionability

Concrete templates exist (Observation/Interpretation/Implication/Next step, mean +/- std, delta vs baseline), but there is no executable guidance for parsing result files, no formula or command for computing deltas, and no example of the required "raw data table" format — pseudocode-level direction rather than fully executable instruction.

3 / 5

Workflow Clarity

A clear, well-sequenced 5-step workflow with an implicit checkpoint ("Flag outliers or suspicious results"), but there is no explicit validation step such as sanity-checking parsed data before analysis; read-only analysis means the destructive-operation cap does not apply.

4 / 5

Progressive Disclosure

Under 50 lines with no need for external references (none exist in the bundle), and content is organized into clearly headed sections — meeting the simple-skill exception for a top score.

5 / 5

Total

17

/

20

Passed

Description

76%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with an explicit what-and-when structure and concrete capability verbs. Its main weakness is trigger coverage: the generic term "compare" invites mis-triggering and natural synonyms for ML result analysis are absent.

Suggestions

Replace or qualify the generic trigger "compare" with domain-specific phrases like "compare model runs", "compare experiment results", or "which config performed best".

Add natural synonyms users would say, such as "experiment results", "metrics", "baseline comparison", or "sweep results", to broaden trigger coverage.

DimensionReasoningScore

Specificity

Lists several concrete actions — "compute statistics, generate comparison tables and insights" — but "compute statistics" and "insights" are generic and coverage has minor gaps (e.g., no mention of visualization or report generation), so it falls just short of the comprehensive 5 anchor.

4 / 5

Completeness

Explicitly answers both: a clear "what" ("Analyze ML experiment results, compute statistics, generate comparison tables and insights") and a "when" with concrete trigger phrases ("Use when user says 'analyze results', 'compare', or needs to interpret experimental data").

5 / 5

Trigger Term Quality

"analyze results" and "interpret experimental data" are natural phrases users would say, but the bare term "compare" is highly generic and common variations/synonyms ("experiment results", "metrics", "seeds", "baseline comparison") are missing.

3 / 5

Distinctiveness Conflict Risk

"ML experiment results" carves a mostly distinct niche, but the standalone trigger "compare" overlaps with diff/review/comparison skills, creating minor conflict risk rather than the minimal risk of a 5.

4 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.