Content
75%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A well-structured, lean persona skill: concrete adversarial techniques with specific probes, an explicit depth-calibration gate, confidence thresholds that serve as validation checkpoints, and a territory section that avoids overlap with sibling reviewers. Uniformly strong with only minor gaps — a vague 'Standard' depth tier and no findings-output format.
Suggestions
Quantify the 'Standard' depth tier (e.g., '1000-3000 words or 5-10 requirements, one or two risk signals') so the middle case is as decidable as Quick and Deep.
Add brief guidance on the findings output format (e.g., one block per finding: premise/assumption quoted, counterargument, consequence, confidence) so results are consistently structured.
State explicitly in the depth-calibration section which techniques map to which tier when signals conflict (e.g., few requirements but a high-stakes domain), since the three tier definitions can overlap.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Nearly every bullet instructs a technique with concrete conditions ("what happens at 10x? At 0.1x?", "High reversal cost + low evidence quality = risky decision") and nothing explains concepts Claude already knows. It sits just below the lean-everywhere anchor 5 because the depth-calibration tier descriptions and a few bullets could be tightened slightly. | 4 / 5 |
Actionability | For an instruction-only skill, the guidance is highly actionable: each of the five techniques has named probes with specific conditions and consequences (falsification test, reversal cost, subtraction test, do-nothing baseline). The minor gap keeping it from 5 is that the 'Standard' depth tier is defined vaguely ("medium document, moderate complexity") while Quick and Deep have quantified thresholds. | 4 / 5 |
Workflow Clarity | The sequence is clear: calibrate depth from explicit size/risk signals, run the technique set for that depth, then apply confidence calibration, with "Below 0.50: Suppress" acting as an explicit validation checkpoint. Anchor 4 fits: most checkpoints are present, but there is no explicit guidance on how to structure or order the findings output itself. | 4 / 5 |
Progressive Disclosure | The skill is a single self-contained file with well-organized, clearly headed sections (depth calibration, five techniques, confidence calibration, territory) and no bundle files to reference; nothing clearly belongs in a separate file. The under-50-lines exception for a score of 5 does not apply at this length, so the good-structure anchor 4 is the best fit. | 4 / 5 |
Total | 16 / 20 Passed |