CtrlK
BlogDocsLog inGet started
Tessl Logo

ablation-planner

Use when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission.

56

Quality

64%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/ablation-planner/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-structured, highly actionable workflow: concrete file sources, a complete Codex prompt spec, feasibility gating, smoke tests, and explicit rules against silent cuts and no-op ablations. Its only weaknesses are minor — a slightly redundant output template and implicit error-recovery after failed smoke tests.

DimensionReasoningScore

Conciseness

The body is lean and assumes Claude's competence — no ML or ablation basics are explained, and rules are stated tersely ("No 'just try it' experiments"). The Step 3 normalization template with generic example rows ("remove module X", lambda values) partially duplicates the format Codex already outputs, which is the minor trimmable instance keeping it at anchor 4 rather than 5.

4 / 5

Actionability

Concrete throughout: exact source files ("idea-stage/docs/research_contract.md", "EXPERIMENT_LOG.md"), a copy-paste-ready MCP invocation with model and reasoning config, a structured output template, and naming conventions ("ablation-no-module-X"). Minor gaps — e.g., "update findings.md with insights" gives no format or location — hold it at anchor 4 rather than fully-executable anchor 5.

4 / 5

Workflow Clarity

A clear 5-step sequence with real checkpoints: feasibility review before running, "Smoke test each ablation before full run", and "don't silently drop ablations" with Codex re-prioritization. The smoke-test failure path (fix and retry) is only implicit, so it sits at anchor 4 ('most checkpoints present; minor validation gaps') rather than the explicit-feedback-loop anchor 5.

4 / 5

Progressive Disclosure

A single self-contained file with well-organized sections (When to Use, Workflow, Rules) and no bundle files or nested references. Nothing here belongs in a separate file — the whole content is one cohesive workflow — so navigation is trivial and the structure matches the appropriate-split intent of anchor 5.

5 / 5

Total

17

/

20

Passed

Description

51%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description has an explicit, distinctive trigger condition but completely omits the 'what' — it never says the skill designs or plans ablation studies. Adding a leading capability clause would lift both completeness and specificity substantially.

Suggestions

Lead with a 'what' clause in third person, e.g., 'Designs and prioritizes ablation studies for ML papers. Use when main results pass result-to-claim...' — this fixes the completeness gap directly.

State 1-2 concrete actions the skill performs (e.g., 'generates an ablation plan with priorities and compute estimates, reviews feasibility, and runs experiments') to raise specificity from a domain-only description.

Add a few natural trigger synonyms such as 'ablation experiments', 'reviewer-requested ablations', or 'ML paper submission' so users phrasing the need differently still match.

DimensionReasoningScore

Specificity

The description names the domain ("ablation studies are needed for paper submission") but never states what the skill actually does — no concrete action verb like 'designs' or 'plans' appears. This matches the anchor 'Names the domain but actions are minimal or generic'; it is not a 3 because no concrete skill action is listed.

2 / 5

Completeness

The 'when' is explicit ("Use when main results pass result-to-claim... and ablation studies are needed"), but the 'what' is absent — the skill's actual capability is never stated, only implied by the trigger condition. Per the guideline to score only what is explicitly stated, this matches anchor 2 ('only when present without what') rather than anchor 3.

2 / 5

Trigger Term Quality

Includes natural terms users would say: "ablation studies", "paper submission", and the concrete gate "result-to-claim (claim_supported=yes or partial)". Good coverage but missing common variations like 'reviewer-requested ablations', 'ML paper', or 'ablation experiments', so it falls short of the comprehensive-synonym anchor 5.

4 / 5

Distinctiveness Conflict Risk

A clear niche (ablation study planning) gated on a specific upstream workflow state ("claim_supported=yes or partial" from result-to-claim). Such specific triggers make it highly distinguishable from other skills with minimal conflict risk, matching anchor 5.

5 / 5

Total

13

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.