Content
86%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A tight, high-signal instruction skill: a lean body with no over-explanation, a mostly concrete workflow anchored by an evidence ladder and an explicit prove-it-or-mark-it-unproven validation step, and clean single-file organization appropriate to its size. The only soft spot is a handful of abstract hints in step 3 and the lack of a fix-and-retry loop that keeps actionability and workflow clarity just below top marks.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The ~45-line body is lean and assumes Claude's competence throughout: it never explains what a diff, grep, or teardown is, and every line carries instruction weight (e.g. 'Listing the callers is not the job. The agent can grep those in a second. The job is the breakage grep won't show you.'). This matches anchor 5 ('Lean and efficient; assumes Claude's competence; every token earns its place'); not 4 because there is no padded or over-explanatory passage that would need trimming — the brief framing sentences ('A blast-radius writeup that sounds right is worthless') justify a rule rather than padding. | 5 / 5 |
Actionability | For an instruction-only skill the guidance is largely concrete: a 5-level evidence ladder ('You ran it. A script or test that calls the real code and fails loud if you're wrong'), a recipe for step 4 ('one small script that imports the same library the app ships and calls the exact function you're worried about'), and a specific output checklist. Not 5 because a few directions stay abstract (e.g. 'Work out when things run: microtasks, unmount and teardown, Solid versus React' — hints rather than executable steps, and oddly specific to Solid/React); not 3 because nothing is pseudocode-level and the key operations (read diff, find safety fact, write and run a proof script, paste output) are directly executable. | 4 / 5 |
Workflow Clarity | The 6 'Steps' are clearly sequenced (read change → find the safety fact → look where grep stops → assess risks → prove the fact → fan out for big changes) with an explicit validation checkpoint ('Write a script or test that runs the real code, run it, and paste what happened') and a failure path ('If you couldn't prove it, write unproven'). This fits anchor 4 ('Clear sequence with most checkpoints present; minor validation gaps'). Not 5 because there is no fix-and-retry feedback loop and the evidence ladder, the steps, and the handback checklist are three parallel structures the reader must merge themselves; not 3 because validation is explicitly present, not missing or implicit. | 4 / 5 |
Progressive Disclosure | No bundle files exist (no references/, scripts/, or assets/), and the self-contained body is under 50 lines with well-organized sections ('Don't trust your own writeup', 'How sure are you', 'Steps', 'What to hand back') — squarely the simple-skill case the rubric says can score 5 on organization alone. All cross-references are to sibling skills (`how`, `why`, `unslop`, `arena`), not nested files, so there is no reference depth or navigation problem. Not 4 because nothing that belongs in a separate file is inlined and no organization gaps exist. | 5 / 5 |
Total | 18 / 20 Passed |