Content
81%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A well-engineered runbook: correct sequencing, explicit validation and feedback loops at every risky point, cost transparency before the expensive phase, and honest failure-mode guidance grounded in recorded pilot runs rather than generic advice. The two systematic weaknesses are rhetorical over-justification in a few passages and sub-agent dispatch instructions that are described rather than given as a verbatim prompt template.
Suggestions
Tighten the justificatory prose in the Step 2 stall-recovery paragraph and the Step 3 'never fabricate a token count' passage to imperative statements — the rules are valuable but the multi-sentence rationales could be roughly halved without losing the pilot-run citations.
Add a verbatim sub-agent prompt template for one (variant, task) cell in Step 2 (the way model-right-sizer-dryrun instructs it), so the 96 dispatches are copy-paste consistent rather than assembled from a description.
Consider moving the Step 2 stall-diagnosis/recovery procedure and the Step 3 dispatch-overhead normalization guidance into a linked reference file, keeping SKILL.md to the runbook proper.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Nearly all of the body is project-specific knowledge Claude cannot already know (pilot-run artifacts, the runtime's 20-dispatch concurrency cap, the ~5-20x whole-dispatch overhead measurement artifact, scope cost math), and the two heredoc scripts are lean. It is not a 5 because several passages carry justificatory padding that could be trimmed — "This is load-bearing, not a style preference", "visible, not assumed away", and the multi-sentence rationales after "Never fabricate a token count" and "a made-up accuracy figure is worse than a visibly incomplete one, because it's indistinguishable from a real one in the report". It stays above anchor 3 because none of it explains concepts Claude already knows — every long passage encodes a pilot-learned failure mode. | 4 / 5 |
Actionability | Two copy-paste-ready executable blocks (the variant-generation python heredoc and the metrics-computation python heredoc), an exact validation command (`uv run --no-project --with jsonschema scripts/validate_blueprint.py -`), exact output paths (`$SCRATCH/results/composition/<variant-name>/<task-id>.json`), and a per-row recipe for the accuracy sweep. It misses anchor 5 only because the sub-agent dispatch instructions in Step 2 are described precisely but never given as a verbatim prompt template, and `SCRATCH=<a scratch directory — never a path inside the repo checkout>` is a placeholder rather than a concrete default — minor gaps, not the missing-steps pattern of anchor 3. | 4 / 5 |
Workflow Clarity | Five clearly sequenced steps with explicit validation and feedback loops throughout: a pre-flight scope confirmation with computed cost counts before the expensive phase; schema validation of every blueprint with a re-emit-once retry loop ("if it doesn't validate, ask that one sub-agent to re-emit once, quoting the validator's error"); a byte-identical check of the all-four variant delegated to `tests/model_right_sizer/test_ablation_layers.py`; a stall-diagnosis and recompute-from-filesystem recovery procedure for batch dispatches; and an explicit "don't infer 'still running' purely from the absence of a failure" checkpoint. This matches the anchor-5 pattern of validate → fix → retry with recovery loops for a batch operation. | 5 / 5 |
Progressive Disclosure | No bundle files exist under references/, scripts/, or assets/ — the body is the whole skill, with all methodology pushed one level out to clearly signaled repo files (`../../eval/ablation/DESIGN.md` — "read that file first; this SKILL.md is the runbook, not the methodology essay" — plus layers.py, benchmark_tasks.json, metrics.py, and a Related section). The body self-describes as a runbook and appropriately defers detail rather than inlining it. It misses anchor 5 because some long inline digressions (the ~40-line stall-recovery procedure and the dispatch-overhead normalization discussion) read as reference material that could be split into their own linked file, and none of the referenced external files are part of this bundle, so a reader outside the parent repo cannot follow them. | 4 / 5 |
Total | 17 / 20 Passed |