CtrlK
BlogDocsLog inGet started
Tessl Logo

results-analysis

This skill should be used when the user asks to "analyze experimental results", "run strict statistical analysis", "compare model performance", "generate scientific figures", "check significance", "do ablation analysis", or mentions interpreting experiment data with rigorous statistics and visualization. It focuses on strict analysis bundles, not Results-section prose.

64

Quality

78%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/results-analysis/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

73%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, disciplined instruction skill with an explicit workflow, validation gates, a QA checklist, and genuine one-level-deep references. The main costs are redundant restatements of audit-mode and blocker rules, and dangling references to a non-existent examples/ directory.

Suggestions

Consolidate the three read-only-audit-mode passages (intro paragraph, 'Non-negotiable quality bar' exception, and the 'Read-only audit mode' section) into one authoritative section and reference it elsewhere in one line.

Merge the overlapping blocker content between step 1 validation, the failure-mode policy, and the final QA gate so each constraint (subject-x-task repeated measures, missing seeds, contradictory interpretation) is stated once.

Either add the examples/ files listed under 'Example files' or remove that section; also consider inlining or relocating the ../research-ideation/references/research-contract.md cross-skill reference so all paths resolve inside the skill bundle.

DimensionReasoningScore

Conciseness

No padding with concepts Claude already knows (no explanation of p-values, CIs, or plotting libraries), but the same rules are stated repeatedly: read-only audit mode is specified in the intro, again in a dedicated section, and echoed in the quality-bar exception; the subject-x-task repeated-measure blocker appears in step 1 and again in the failure-mode policy; quarantine rules appear twice. This is more than the minor trimming the 4 anchor describes, though never vague or padded enough to fall to 2.

3 / 5

Actionability

Concrete, instruction-level guidance throughout: exact output filenames and directory tree, an explicit per-figure requirement list (purpose, plotted variables, error-bar meaning, caption requirements, interpretation checklist), a copy-paste-ready claim-candidate markdown template, and specific statistical deliverables (mean +/- std, 95% CI, effect sizes, multiple-comparison handling). Falls short of 5 because the statistical execution itself is delegated to references and some directives remain abstract ("use non-parametric fallback when assumptions fail" without naming the fallback tests).

4 / 5

Workflow Clarity

A clearly sequenced 6-step workflow with explicit validation checkpoints: step 1 validates artifacts and unit of analysis with a hard stop ("If the comparison is not statistically valid, say so before continuing"), step 2 locks comparison questions before running statistics, and the Final QA gate is an explicit do-not-finish-until checklist. Feedback loops for error recovery are present via the failure-mode policy and the quarantine rule for contradictory statistics files, matching the 5 anchor rather than the 4 anchor with minor validation gaps.

5 / 5

Progressive Disclosure

The body is an overview with six real, on-topic reference files (all verified present in references/) listed with one-line purposes under "Load only what is needed" and signaled in-context at the relevant workflow steps — solid one-level-deep structure. However, the "Example files" section lists three files (examples/example-analysis-report.md, example-stats-appendix.md, example-figure-catalog.md) that do not exist in the bundle, and one reference path points outside the skill directory (../research-ideation/references/research-contract.md), which is unverifiable here — keeping it below the 5 anchor but above the 3 anchor, since the defect is confined to peripheral paths.

4 / 5

Total

16

/

20

Passed

Description

82%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with explicit, natural trigger phrases and a clear niche, plus an explicit scope exclusion. Its main weakness is that the 'what' is carried by quoted user utterances and a jargon-bearing scope sentence rather than a direct capability statement.

DimensionReasoningScore

Specificity

Concrete actions are named via quoted triggers — "run strict statistical analysis", "compare model performance", "generate scientific figures", "check significance", "do ablation analysis" — plus a scope statement ("It focuses on strict analysis bundles, not Results-section prose"). Falls below the 5 anchor because the capabilities are framed as user utterances rather than an explicit capability list, and the bundle outputs (analysis-report.md, stats-appendix.md, figures) are not named.

4 / 5

Completeness

Explicit 'when' guidance ("This skill should be used when the user asks to...") with concrete trigger phrases, and a 'what' present via the embedded action phrases and "It focuses on strict analysis bundles, not Results-section prose". Not a 5 because the standalone 'what' statement is indirect and uses internal jargon ("strict analysis bundles") rather than an explicit capability sentence like 'Runs rigorous statistics and generates scientific figures'.

4 / 5

Trigger Term Quality

Comprehensive natural-term coverage with synonyms: "analyze experimental results", "statistical analysis", "compare model performance", "scientific figures", "check significance", "ablation analysis", "interpreting experiment data", "statistics and visualization". These are phrases a user would naturally say; only the self-coined qualifier "strict" is slightly unnatural, and no meaningful natural synonym is missing.

5 / 5

Distinctiveness Conflict Risk

Clear niche (rigorous experimental-results statistics and figure generation) with a boundary cue ("not Results-section prose") that separates it from manuscript-writing skills. Minor overlap risk remains with generic data-visualization or plotting skills ("generate scientific figures") and general data-analysis skills; the neighboring-skill exclusions named only in the body would sharpen this.

4 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
Galaxy-Dawn/claude-scholar
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.