CtrlK
BlogDocsLog inGet started
Tessl Logo

experiment-plan

Turn a refined research proposal or method idea into a detailed, claim-driven experiment roadmap. Use after `research-refine`, or when the user asks for a detailed experiment plan, ablation matrix, evaluation protocol, run order, compute budget, or paper-ready validation that supports the core problem, novelty, simplicity, and any LLM / VLM / Diffusion / RL-based contribution.

72

Quality

90%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is a tightly structured, claim-driven experiment-planning skill with explicit phases, decision gates, and copy-ready output templates, assuming Claude's expertise rather than over-explaining. It is action-oriented and well-sequenced with validation checkpoints throughout.

Suggestions

Tighten the Key Rules section by removing items that merely restate guidance already embedded in Phase 2–4 of the Workflow (e.g. 'Prefer strong baselines' duplicates the Phase 2 baseline guidance).

Consider moving the EXPERIMENT_PLAN.md and EXPERIMENT_TRACKER.md markdown templates into a references/ file so SKILL.md reads as a leaner overview with clearly signaled one-level-deep references.

DimensionReasoningScore

Conciseness

The body is well-structured and assumes Claude's competence (no basic explanations of what experiments or ablations are), though some Key Rules restate guidance already given in the Workflow phases, leaving minor trim opportunities.

4 / 5

Actionability

Provides fully executable output templates (EXPERIMENT_PLAN.md and EXPERIMENT_TRACKER.md with complete markdown structures), explicit constants, per-block specification fields, and a ready-to-present summary format covering the common cases.

5 / 5

Workflow Clarity

A clear six-phase sequence (Phase 0–5) with milestone structure, explicit stop/go decision gates, success/failure interpretation criteria, must-run vs nice-to-have separation, and a Final Checklist provides explicit validation checkpoints and feedback loops.

5 / 5

Progressive Disclosure

No bundle directories (references/, scripts/, assets/) exist, so all content is inline; the shared protocol references are one-level-deep and clearly signaled, and sections are well-organized, though the full output templates could alternatively live in separate reference files.

4 / 5

Total

18

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is precise, action-oriented, and clearly states both the skill's purpose and its trigger conditions, with concrete domain-specific terms a researcher would naturally use. It is third-person, well-scoped, and distinct from neighboring skills in the pipeline.

DimensionReasoningScore

Specificity

Lists multiple concrete actions and artifacts ("claim-driven experiment roadmap", "ablation matrix", "evaluation protocol", "run order", "compute budget", "paper-ready validation"), giving comprehensive coverage of what the skill produces.

5 / 5

Completeness

Explicitly answers both what ("Turn a refined research proposal...into a detailed, claim-driven experiment roadmap") and when ("Use after research-refine, or when the user asks for...") with concrete trigger phrases.

5 / 5

Trigger Term Quality

Strong natural trigger phrases ("detailed experiment plan", "ablation matrix", "evaluation protocol", "run order", "compute budget") with good synonym coverage, but a few natural variations a researcher might say (e.g. "baseline comparison") are absent.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (post-refinement experiment planning) and is explicitly positioned against companion skills (research-refine, run-experiment), minimizing conflict risk.

5 / 5

Total

19

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

relative_links

Relative link issues: 3 suspicious

Warning

Total

14

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.