Content
61%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The body delivers a clear iterative debug-test loop and sensible guidance on real-vs-mock testing, and it is reasonably token-efficient. Its weaknesses are placeholder-only code examples, a missing debug-script example, broken step numbering (3 jumps to 5), and an unreconciled contradiction about when to run other tests.
Suggestions
Fix the step numbering (the list jumps from 3 to 5) and reconcile the contradiction between 'ignore other tests until you have a working example' and step 7 'Make sure other tests pass'.
Include one minimal executable debug-script example instead of only placeholder paths like './path/to/cli.sh', and show a concrete API call sample for the server case.
Remove the fake <system-reminder> injections and the duplicated 'ignore other tests' instruction; fold them into a single explicit step to tighten conciseness.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is lean with no padding or explanations of concepts Claude already knows, and code examples are compact. Minor tightening is possible: the closing note "Do NOT get in a loop where you just keep running other tests" repeats the earlier "ignore any existing tests until you have a working example", fitting 'Efficient; minor instances of over-explanation that could be trimmed' rather than the fully lean anchor 5. | 4 / 5 |
Actionability | Concrete elements exist (bash/Python/Node snippets for CLI calls, loop steps), but they use placeholder paths like "./path/to/cli.sh" and the API section only says "Call to the server using scripting language of choice" with no example. There is no example of an actual debug script — the core deliverable — matching 'Some concrete guidance but incomplete; pseudocode instead of executable code; missing key details'. | 3 / 5 |
Workflow Clarity | A real feedback loop is specified (add logs → run script → analyze → update, repeated until fixed), but the numbered list skips from step 3 to step 5, and instructions contradict each other: "ignore any existing tests until you have a working example" vs. step 7 "Make sure other tests pass", with no reconciliation of when the switch happens. This fits 'Steps listed but validation gaps; sequence present but checkpoints missing or implicit' rather than the mostly-checkpointed anchor 4. | 3 / 5 |
Progressive Disclosure | The skill is short (<50 lines) with no bundle files, and sections are reasonably organized with one clearly signaled external reference ("read the .claude/skills/webapp-testing/SKILL.md"). Under the simple-skill guidance this could score 5, but the unnumbered-to-body structure (the <required> block restating steps that later sections cover) leaves minor organization gaps, placing it at anchor 4. | 4 / 5 |
Total | 14 / 20 Passed |