CtrlK
BlogDocsLog inGet started
Tessl Logo

aris-experiment-plan

Turn a refined research proposal or method idea into a detailed, claim-driven experiment roadmap. Use after `aris-research-refine`, or when the user asks for a detailed experiment plan, ablation matrix, evaluation protocol, run order, compute budget, or paper-ready validation that supports the core problem, novelty, simplicity, and any LLM / VLM / Diffusion / RL-based contribution.

72

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

SKILL.md
Quality
Evals
Security

Quality

Content

77%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-sequenced, highly actionable planning skill with concrete templates, constants, and verification checkpoints. Its main weaknesses are mild redundancy between the Key Rules and the phase descriptions, and a monolithic structure whose large output templates could be factored into reference files.

Suggestions

Collapse the 'Key Rules' section into the phases it duplicates (simplicity, frontier, strong baselines, must-run vs nice-to-have) to remove redundancy, or restate only the non-obvious rules.

Move the EXPERIMENT_PLAN.md and EXPERIMENT_TRACKER.md template blocks into reference files under references/ and link to them, so the main body stays a lean overview.

Add an explicit validate/retry loop note for the run order (e.g. re-evaluate claims after the sanity stage) to strengthen the already-present decision gates.

DimensionReasoningScore

Conciseness

Mostly efficient and assumes Claude's competence (no preamble on what ablations or seeds are), but the 'Key Rules' section repeats guidance already given in the phases (e.g. defending simplicity, preferring strong baselines) and could be tightened. Not 3 because of this redundancy; not 1 because it is not padded with concepts Claude already knows.

2 / 3

Actionability

Provides fully specified, copy-paste-ready output templates with exact field lists per experiment block, fixed constants (MAX_PRIMARY_CLAIMS=2, DEFAULT_SEEDS=3), and concrete milestone/summary structures. Not 2 because guidance is executable and complete, not pseudocode or vague direction.

3 / 3

Workflow Clarity

Sequences a clear Phase 0-5 process with explicit stop/go decision gates per milestone and a final 'Final Checklist' verification step. Not 2 because validation checkpoints are present rather than merely implicit.

3 / 3

Progressive Disclosure

Well-organized into sections, but the 237-line body is monolithic with large inline template blocks (EXPERIMENT_PLAN.md and EXPERIMENT_TRACKER.md structures) that could live in separate reference files. Not 3 because it exceeds the simple-skill carve-out without splitting detail into one-level-deep references; not 1 because sections are clearly labeled and navigable.

2 / 3

Total

10

/

12

Passed

Description

100%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, third-person description that states concrete deliverables, gives an explicit usage trigger, and names the natural terms a researcher would request. It is clearly distinguishable from sibling research skills and does not over-claim.

DimensionReasoningScore

Specificity

Lists multiple concrete deliverables — 'experiment roadmap', 'ablation matrix', 'evaluation protocol', 'run order', 'compute budget', 'paper-ready validation' — rather than vague language. Not 2 because it goes beyond naming the domain to enumerating specific outputs.

3 / 3

Completeness

Explicitly answers both what ('Turn a refined research proposal... into a detailed, claim-driven experiment roadmap') and when ('Use after aris-research-refine, or when the user asks for...'). Not 2 because the trigger is explicit, not merely implied.

3 / 3

Trigger Term Quality

Uses natural terms a researcher would actually say — 'experiment plan', 'ablation matrix', 'evaluation protocol', 'run order', 'compute budget'. Not 2 because it covers the common phrasings a user would request, not just one keyword.

3 / 3

Distinctiveness Conflict Risk

Occupies a clear niche — paper-oriented experiment planning for ML/LLM/VLM/Diffusion/RL research — with triggers tied to a specific skill pipeline, making conflicts unlikely. Not 2 because it is far more specific than a generic 'research' skill.

3 / 3

Total

12

/

12

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

Total

15

/

16

Passed

Repository
OpenLAIR/dr-claw
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.