CtrlK
BlogDocsLog inGet started
Tessl Logo

ablation-planner

Use when main results pass result-to-claim (`claim_supported = yes` or `partial`) and ablation studies are needed for paper submission. A secondary Codex agent designs ablations from a reviewer's perspective; the local executor reviews feasibility and implements.

65

Quality

78%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/skills-codex/ablation-planner/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured orchestration skill: the reviewer/executor division of labor is enforced in both workflow and Rules, the design prompt and plan schema are fully specified, and feasibility checks plus smoke tests gate execution. The main gaps are the absence of an explicit failure-recovery loop after smoke tests and no example command for the implement-and-run phase.

Suggestions

Add one line to Step 5 defining the smoke-test failure path (e.g., 'If a smoke test fails, fix the config and re-run the smoke test before any full run') to close the workflow_clarity validation gap.

Include a minimal example of an ablation config or launch command (e.g., a config diff naming convention like `ablation-no-module-X`) to make the implement phase copy-paste ready.

Trim the duplication between the Step 2 prompt's output specification and the Step 3 template — reference one from the other — to tighten conciseness.

DimensionReasoningScore

Conciseness

The body is lean and assumes domain competence — it never explains what an ablation is, and steps are given as terse imperatives ('Smoke test each ablation before the full run'). It is not level 5 because there is mild redundancy: the 'When to Use' bullets restate the frontmatter description, and the Step 3 markdown template partially duplicates the structure already specified in the Step 2 prompt. It is comfortably above level 3, whose anchor includes genuinely unnecessary explanation.

4 / 5

Actionability

Concrete, executable guidance dominates: a complete spawn_agent block with model and reasoning_effort, a full reviewer-prompt template with named output fields, an exact normalized-plan markdown schema, named file sources (`idea-stage/docs/research_contract.md`, `EXPERIMENT_LOG.md`), and a fallback path ('If delegation is unavailable, generate the same plan locally and mark it `[pending external review]`'). It falls short of anchor 5's copy-paste-ready coverage because Step 5's implement-and-run phase gives no example command or config skeleton, leaving the actual execution mechanics to be inferred.

4 / 5

Workflow Clarity

A clear five-step sequence with real checkpoints: Step 4 is an explicit pre-run feasibility review (compute budget, code vs config changes, parallelism, proposed cuts with re-prioritization) and Step 5 mandates smoke tests before full runs, satisfying the validation expectation for batch experiment operations. It does not reach anchor 5 because there is no error-recovery loop spelled out for failed smoke tests or contradicting results — the fix-and-retry feedback path present in the anchor-5 example is only implied.

4 / 5

Progressive Disclosure

The file is well-sectioned (When to Use, Workflow steps 1-5, Rules) with no buried references and no nested lookups — there are no bundle files at all, so everything lives in one navigable document. It scores 4 rather than 5 because the skill exceeds the short-skill threshold where section organization alone earns a 5, and the ~60-line design-prompt plus plan-schema block could arguably be split out of the main workflow if the skill grows. It is well above anchor 3, which requires inlined content that should be separate or unclear signaling.

4 / 5

Total

16

/

20

Passed

Description

82%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with an explicit, well-parameterized 'Use when' trigger and a clear third-person statement of the division of labor between the reviewer agent and local executor. The only weakness is that the 'what' stays at the action-summary level rather than naming the concrete outputs (component ablations, sensitivity sweeps, comparisons) the skill produces.

DimensionReasoningScore

Specificity

The description names the domain ('ablation studies', 'paper submission') and two concrete actions — 'A secondary Codex agent designs ablations from a reviewer's perspective; the local executor reviews feasibility and implements' — but does not enumerate what the ablation planning actually produces (component removal, hyperparameter sweeps, design-choice comparisons). This matches anchor 3 ('names domain and 1-2 concrete actions, but not comprehensive'); it is not level 4 because it lacks the several specific actions listed there, and clearly above level 2's generic minimal actions.

3 / 5

Completeness

Both parts are explicit: the trigger is 'Use when main results pass result-to-claim (`claim_supported = yes` or `partial`) and ablation studies are needed for paper submission', and the 'what' states the two-agent design/implement split. This clearly matches anchor 5 ('explicitly answers both what AND when with concrete trigger phrases'), not 4, because the 'when' clause is fully specific with concrete trigger values rather than merely adequate.

5 / 5

Trigger Term Quality

Natural trigger phrases are present — 'ablation studies', 'ablation studies are needed for paper submission', 'paper submission' — plus concrete pipeline triggers ('result-to-claim', 'claim_supported = yes or partial'). A few natural variations users might say are missing (e.g., 'reviewer requests ablations', 'ML paper experiments'), so it sits at anchor 4 rather than 5's comprehensive synonym coverage, but well above anchor 3's sparser keyword sets.

4 / 5

Distinctiveness Conflict Risk

It occupies a clear niche — ablation planning for ML paper submission gated on a specific upstream check — with distinct trigger terms ('result-to-claim', 'claim_supported', 'paper submission') unlikely to collide with other skills, matching anchor 5's 'clear niche with distinct triggers; minimal conflict risk'. Voice is consistently third person, so no specificity penalty applies.

5 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

Total

15

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.