Content
62%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The body excels at workflow clarity — the RED/GREEN/REFACTOR and Prove-It loops are explicitly sequenced with validation checkpoints, feedback loops, and a completion checklist — and its guidance is largely concrete with real code examples. Its two weaknesses are noticeable verbosity from re-teaching standard testing doctrine Claude already knows, and a monolithic structure whose two external references are vaguely signaled and not backed by any bundle files.
Suggestions
Cut or drastically compress sections restating knowledge Claude already has (Arrange-Act-Assert, one-assertion-per-concept, good/bad test naming, mocks-vs-fakes, the Common Rationalizations table) to a few lines of policy each.
Move the Browser Testing with DevTools section into a reference file (e.g., references/browser-testing.md) and link to it, keeping only the security-boundary warning inline.
Fix the dangling references: link browser-testing-with-devtools and testing-patterns.md with clear, valid paths inside the skill's own references/ directory so navigation is unambiguous.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The ~400-line body re-teaches testing knowledge Claude already has at length: the Arrange-Act-Assert pattern, one-assertion-per-concept, descriptive test naming with good/bad examples, the mocks-vs-stubs-vs-fakes preference order, the test pyramid, a "Common Rationalizations" motivational table, and the "Beyonce Rule" quip — several padded sections of standard doctrine, matching "noticeably verbose; several unnecessary explanations or padded sections". It is more purposeful than the heavily padded level 1 (the Discover-the-Stack-First and Prove-It sections add genuine workflow instruction Claude might not follow unprompted), so it does not fall to 1; but the standard-testing-tutorial material keeps it clearly below the "minor instances" level 4. | 2 / 5 |
Actionability | Guidance is mostly concrete and executable: runnable TypeScript test examples for each cycle step ("expect(task.status).toBe('pending')"), a worked bug-reproduction example, a concrete stack-discovery checklist ("package.json, pom.xml/build.gradle, pyproject.toml...") with named wrapper commands ("./gradlew", "make test"), a test-size decision guide, and an end-of-work verification checklist. Minor gaps keep it from 5: the code examples are illustrative fragments (taskService and db are never defined), and no commands are copy-paste ready since the skill deliberately defers to the repo's own tooling. | 4 / 5 |
Workflow Clarity | The core loop is an explicit, clearly sequenced workflow with validation checkpoints at every step — RED requires the test to FAIL ("A test that passes immediately proves nothing"), GREEN requires it to pass, REFACTOR requires "Run tests after every refactor step to confirm nothing broke" — and the Prove-It pattern adds a full feedback loop ending in "Run full test suite (no regressions)". A closing checklist ("Every new behavior has a corresponding test", "Bug fixes include a reproduction test that failed before the fix") matches the top anchor: clear sequence, explicit validation, feedback loops, and checklists. | 5 / 5 |
Progressive Disclosure | Section structure is good, but the skill is monolithic: no references/ bundle exists, yet the body inlines ~40 lines of browser/DevTools testing that clearly belong in a separate file, and it points to two external resources — the bare name "browser-testing-with-devtools" and a path outside the skill directory ("../../references/testing-patterns.md") — that are neither clearly signaled nor present as bundle files. This fits "some structure but could be better organized; references present but not clearly signaled; content that should be separate is inline"; it is well above the unstructured level 2 thanks to clean headers and a genuinely useful overview role. | 3 / 5 |
Total | 14 / 20 Passed |