CtrlK
BlogDocsLog inGet started
Tessl Logo

results-analysis

This skill should be used when the user asks to "analyze experimental results", "run strict statistical analysis", "compare model performance", "generate scientific figures", "check significance", "do ablation analysis", or mentions interpreting experiment data with rigorous statistics and visualization. It focuses on strict analysis bundles, not Results-section prose.

64

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/results-analysis/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

73%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-sequenced, validation-rich workflow with concrete output artifacts and a strong QA gate. Its weaknesses are moderate internal redundancy and several referenced paths (examples/ and a cross-skill contract) that do not exist on disk.

Suggestions

Deduplicate the read-only audit-mode description (currently in both the intro and its own section) and consolidate the inline reference citations with the "Reference files" list to remove repetition.

Create the referenced examples/ files (example-analysis-report.md, example-stats-appendix.md, example-figure-catalog.md) or remove the "Example files" section, since those paths do not exist.

Fix or remove the dangling cross-skill reference ../research-ideation/references/research-contract.md, which is not present in the bundle.

DimensionReasoningScore

Conciseness

The body is operational and avoids basic-concept padding, but contains noticeable redundancy: read-only audit mode is described both in the intro and again in a dedicated section, and the reference files are cited inline then re-listed in a "Reference files" section, so it could be tightened past the efficient anchor 4.

3 / 5

Actionability

Concrete, specific guidance throughout — exact output filenames (analysis-report.md, stats-appendix.md, figure-catalog.md), an exact statistics package (mean ± std, 95% CI, effect sizes, multiple-comparison handling), a claim-candidate template, and a QA checklist — but executable statistical/figure-generation code is deferred to reference files, leaving minor gaps versus copy-paste-ready anchor 5.

4 / 5

Workflow Clarity

A clear six-step Standard workflow is sequenced with explicit validation checkpoints ("If the comparison is not statistically valid, say so before continuing"), a final QA-gate checklist ("Do not finish until all are true"), and error-recovery feedback loops (quarantine policy, failure-mode policy), matching the anchor 5.

5 / 5

Progressive Disclosure

Good one-level-deep structure: six real reference files are cited inline and re-listed with one-line descriptions, but the "Example files" section points to a non-existent examples/ directory and the cross-skill reference ../research-ideation/references/research-contract.md is also missing, which is an organization gap below anchor 5.

4 / 5

Total

16

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description provides strong, explicit trigger guidance and a clear niche boundary, with multiple specific natural phrases. Its main weakness is that the skill's own capabilities are conveyed indirectly through user-trigger phrasing and a scope statement rather than a direct, concrete action list.

DimensionReasoningScore

Specificity

Lists several specific actions ("analyze experimental results", "run strict statistical analysis", "compare model performance", "generate scientific figures", "check significance", "do ablation analysis") with broad coverage, but they are framed as user triggers rather than direct capability statements, leaving a minor gap versus the comprehensive-action anchor.

4 / 5

Completeness

The "when" is explicit and concrete ("should be used when the user asks to ..."), and a "what" is present ("It focuses on strict analysis bundles, not Results-section prose"), but the "what" is a scope statement rather than a crisp enumeration of the skill's own actions, so it sits at anchor 4 rather than 5.

4 / 5

Trigger Term Quality

Six natural quoted trigger phrases ("analyze experimental results", "run strict statistical analysis", "compare model performance", "generate scientific figures", "check significance", "do ablation analysis") give good keyword coverage, but synonyms (e.g. "benchmark", "evaluate") and file-type terms are missing, so it falls short of the comprehensive anchor 5.

4 / 5

Distinctiveness Conflict Risk

Clear niche (strict statistical analysis of experimental results) with distinct triggers and an explicit boundary ("not Results-section prose"), giving minimal conflict risk and matching the clear-niche anchor 5.

5 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
Galaxy-Dawn/claude-scholar
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.