Content
78%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A well-structured rubric skill with excellent actionability: concrete rewrite pairs, exact failure vocabularies, a deterministic severity-to-verdict rule, and a complete output format. The main cost is token weight from repeated ISTQB glossary citations and motivational framing that Claude does not need; the workflow is clearly sequenced but lacks an explicit final verification checkpoint.
Suggestions
Trim the ISTQB glossary quotations (testability, expected result x2, test oracle, equivalence partitioning, boundary value analysis) to a single one-line citation or drop them entirely — Claude already knows these concepts, and the heuristics' own failure signatures carry the load.
Cut the 'shift left' paragraph and aphorisms ('The cheapest defect to fix is the one prevented before it's coded') to one sentence of placement guidance; the workflow section already says when to run the review.
Add an explicit final verification step to the workflow, e.g. 'Before emitting, confirm every non-OK row has a rewrite and the verdict matches the severity table', to close the last workflow-clarity gap.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is mostly tight (failure-signature word lists, compact rewrite tables), but it repeatedly explains concepts Claude already knows: quoted ISTQB definitions of 'testability', 'expected result' (cited twice), 'test oracle', 'equivalence partitioning', and 'boundary value analysis', plus a 'shift left' rationale and aphorisms like 'The cheapest defect to fix is the one prevented before it's coded'. This matches 'mostly efficient but includes some unnecessary explanation or could be tightened'; it is not a 4 because the glossary quoting is a recurring pattern rather than a minor instance. | 3 / 5 |
Actionability | For an instruction-only skill the guidance is fully concrete: three untestable-to-testable rewrite tables ('p95 page-load time on `/dashboard` is at most 2.5s under 100 concurrent users'), exact failure-signature vocabulary ('fast, clean, robust, seamless...'), a deterministic severity table, a verdict rule, and a filled-in output-format findings table. Not a 4: there are no gaps — every step produces a specified artifact with copy-paste-ready structure. | 5 / 5 |
Workflow Clarity | A clear 6-step sequence with built-in checkpoints ('Record every heuristic it fails, not just the first', 'A finding with no proposed replacement is not a finding') and an anti-patterns table that serves as error-recovery guidance. Not a 5: there is no explicit final verification step (e.g., confirming the emitted verdict matches the severity table or that every flagged claim carries a rewrite before emitting), only implicit rules; not a 3 because the sequence and checkpoints are otherwise explicit and the operation is read-only, so the destructive/batch cap does not apply. | 4 / 5 |
Progressive Disclosure | The single bundle file, references/worked-examples.md, exists, is linked with a descriptive signpost ('Three worked reviews - one BLOCK verdict..., one OK verdict..., and one REVIEW verdict... are in [references/worked-examples.md]'), and is exactly one level deep (it links back to SKILL.md, no further nesting). The rubric itself belongs in SKILL.md and the examples are appropriately split out. Not a 4: navigation and placement leave no real gaps. | 5 / 5 |
Total | 17 / 20 Passed |