CtrlK
BlogDocsLog inGet started
Tessl Logo

experiment-plan

Turn a refined research proposal or method idea into a detailed, claim-driven experiment roadmap. Use after `research-refine`, or when the user asks for a detailed experiment plan, ablation matrix, evaluation protocol, run order, compute budget, or paper-ready validation that supports the core problem, novelty, simplicity, and any LLM / VLM / Diffusion / RL-based contribution.

69

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, actionable experiment-planning workflow with clear phases and templates, weakened only by dead shared-reference links and slightly verbose embedded templates. It assumes Claude's intelligence and avoids explaining basics.

Suggestions

Fix or remove the shared-references links (output-versioning.md, output-manifest.md, output-language.md) — they point to files that do not exist in the skill bundle, leaving navigation broken.

Add an explicit validation/checkpoint step after each milestone's outputs are written (e.g. confirm EXPERIMENT_PLAN.md covers every claim and tracker rows match blocks) to strengthen feedback loops for the batch writes.

Tighten the two embedded markdown templates to the essential fields only, or move the full templates into a references file, to improve token efficiency.

DimensionReasoningScore

Conciseness

Lean and assumes Claude's competence with no padding about what experiments are, but the two full markdown templates (EXPERIMENT_PLAN and EXPERIMENT_TRACKER) add length that could be slightly trimmed; efficient but not maximally lean.

4 / 5

Actionability

Concrete block specs (claim tested, datasets, metrics, success criterion, failure interpretation), a milestone run order with cost/risk estimates, and copy-ready output templates provide mostly executable guidance with only minor gaps.

4 / 5

Workflow Clarity

Clear six-phase sequence (0-5) with stop/go decision gates, must-run vs nice-to-have separation, and a final checklist; minor validation gaps in that individual milestones lack an explicit 'verify output' checkpoint despite batch file writes.

4 / 5

Progressive Disclosure

Good section structure and a clear overview, but the only external references (../../shared-references/output-versioning.md, output-manifest.md, output-language.md) do not exist as real files, so the references are dead and not effectively one-level-deep.

3 / 5

Total

15

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, specific, well-triggered description that clearly states both capability and usage conditions while distinguishing itself from related pipeline skills. No vague fluff or over-claims.

DimensionReasoningScore

Specificity

Names the domain and multiple concrete artifacts — 'ablation matrix, evaluation protocol, run order, compute budget, paper-ready validation' — giving comprehensive coverage of what the skill produces.

5 / 5

Completeness

Explicitly states both what it does ('Turn a refined research proposal... into a detailed, claim-driven experiment roadmap') and when to use it ('Use after research-refine, or when the user asks for...').

5 / 5

Trigger Term Quality

Rich natural trigger coverage ('detailed experiment plan, ablation matrix, evaluation protocol, run order, compute budget, paper-ready validation') matching phrases a researcher would actually say.

5 / 5

Distinctiveness Conflict Risk

Occupies a clear niche ('experiment-plan') and is positioned against sibling skills (research-refine, run-experiment, auto-review-loop), minimizing overlap with other skills.

5 / 5

Total

20

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 3 suspicious

Warning

Total

15

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.