Content
77%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A strong, disciplined runbook: explicit costs, gated spend, an independent-evaluation protocol with feedback loops for stall recovery, and honest held-out/overfit reporting make the workflow itself exemplary. Weaknesses are moderate: copy-paste command invocations are missing, and the skill is structurally dependent on DESIGN.md and the eval/tuning scripts that are not shipped inside the bundle, leaving the body's "read that first" instruction unverifiable from the skill alone. Minor redundancy (repeated never-edit rule, duplicate link lists) costs a little conciseness.
Suggestions
Progressive disclosure: ship the load-bearing dependencies inside the skill bundle (e.g. references/DESIGN-excerpt.md, scripts/ for knobs/optimizer/generate_variant) or inline the essential content — currently every ../../eval/tuning/* path lives outside the bundle while the body says 'read that file first' and Step 4 requires restating DESIGN.md's 'Known limitations' verbatim.
Actionability: add copy-paste-ready command lines for the tooling the loop depends on (e.g. the exact generate_variant.py and optimizer invocations for rendering and scoring a neighbor), since only function signatures are currently given.
Conciseness: remove the duplication between Step 4, 'What this skill does NOT do', and 'Related' — the never-edit rule is stated twice and all four eval/tuning files are linked inline and again in Related.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is dense and almost entirely operational — concrete costs ("typically 3–9 real builds per candidate", "sonnet 38,401 / haiku 28,508 tokens"), specific tool functions, and no explanation of concepts Claude already knows. It falls at the 4 anchor ("Efficient; minor instances of over-explanation that could be trimmed") rather than 5 because of real duplication: the "never edits the shipped file" rule is stated in both Step 4 and "What this skill does NOT do", the `../../eval/tuning/` files are each linked inline and again in "Related", and asides like "this SKILL.md is the runbook; the design doc is the honest account" are editorial padding. It is clearly not 3 — no whole section is unnecessary. | 4 / 5 |
Actionability | Guidance is mostly executable with named interfaces throughout: "`knobs.ALL_KNOBS`", "`optimizer.propose_neighbors`", "`optimizer.score_candidate(records)`", "`optimizer.coordinate_ascent_step(current_settings, current_score, knob_name, neighbor_evaluations)`", the concrete tuning/held-out task IDs, and a worked recovery reference for the stall case. It sits at 4 ("Mostly executable guidance; concrete code or commands with minor gaps") rather than 5 because there are no copy-paste-ready command lines — e.g. how to actually invoke `generate_variant.py` or `optimizer` from the CLI is never shown, so the operator must reconstruct invocations from function signatures. | 4 / 5 |
Workflow Clarity | The process is explicitly sequenced (scope confirmation → Step 1 floors → Step 2 search loop → Step 3 held-out check → Step 4 report) with validation checkpoints throughout: a spend-confirmation gate before any builds ("ask before spending real build budget"), a convergence criterion ("if no knob moved, stop — converged"), a held-out overfit check whose failure is elevated to "the headline finding", and an explicit feedback loop for stall recovery ("diagnose via the run's `journal.jsonl` (`started` vs. `result` event counts) and output-file mtimes... recover by recomputing the missing set from the filesystem"). This matches the 5 anchor (explicit validation steps, feedback loops for error recovery, a checklist for the final report); the batch-operation cap does not apply because validation is present. | 5 / 5 |
Progressive Disclosure | Structure and signaling are good — references are one level deep and clearly marked (DESIGN.md, knobs.py, optimizer.py, generate_variant.py, the sibling skill), and the body is well-sectioned. But the score falls to the 3 anchor rather than 4 because the runbook's load-bearing content is deferred to files that are not part of the skill bundle: no `references/`, `scripts/`, or `assets/` exist, and every "[`../../eval/tuning/...`]" path resolves outside the bundle (unresolvable in this layout), while the body insists "**read that file first**" and Step 4 requires restating "Every limitation named in DESIGN.md's 'Known limitations' section" — content the bundle does not ship. The duplicated "Related" link list adds organization noise but is secondary to the unshipped-dependency gap; it is not 2 because the SKILL.md itself is a coherent, well-organized runbook rather than an inlined dump. | 3 / 5 |
Total | 16 / 20 Passed |