Content
77%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
An exceptionally actionable, well-sequenced workflow skill: exact commands, prompts, templates, verdict precedence, and validation gates throughout. Its weaknesses are length — ~780 lines with the qa.level/testing.strict policy explained twice and rationale commentary that could be trimmed (including one editing artifact at line 129) — and a monolithic single-file structure that inlines content a reference file should carry.
Suggestions
State the qa.level/testing.strict resolution rules once (Phase 1 or Phase 3) and cross-reference from the other phase instead of re-explaining them at length.
Move stable, rarely-changed material — the evidence-gate table, verdict definitions, and the tech-debt register row format — into a reference file (e.g. references/evidence-gates.md) to cut the body's length and duplication.
Fix the editing artifact at line 128–129 where "that should be in localization files" dangles after the bolded code-root clause, restoring the original 'Grep for player-facing strings that should be in localization files' instruction.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is dense and project-specific rather than padded with concepts Claude already knows, but the qa.level/testing.strict policy is expounded at length twice (Phase 1 and Phase 3), run-result rules repeat across phases, and design-rationale commentary plus an editing artifact (the dangling "that should be in localization files" at line 129) leave more than minor trimming. Anchor 3 rather than 4 because the tightening opportunities are substantial, but well above the verbosity of anchor 2. | 3 / 5 |
Actionability | Guidance is fully concrete: exact Grep patterns and output modes, exact AskUserQuestion prompts and option labels, exact report/table/checkpoint templates, exact commands (e.g. `bash .claude/scripts/story-status.sh`) and exact tech-debt row formats. As an instruction-only skill this is copy-paste-ready guidance covering the common cases, matching anchor 5. | 5 / 5 |
Workflow Clarity | Eight explicitly sequenced phases with validation before every write (report presented before any file edit, explicit user approval in Phase 7), feedback loops on failure (BLOCKED does not auto-proceed, "supply it and re-run the gate"), and explicit verdict precedence with definitions — a textbook match for anchor 5. The corrupted hardcoded-strings bullet muddies one check but does not degrade the sequence or validation structure. | 5 / 5 |
Progressive Disclosure | Sections are well-organized and external doc references are clearly signaled one level deep, but the skill is a single ~780-line file with no bundle at all — evidence-gate tables, the config-resolution policy exposition, and the tech-debt row format duplicated from another skill are inlined where separate reference files belong. Anchor 3 rather than 4 because the volume of inlinable content is substantial; not 2 because structure and reference signaling are genuinely good. | 3 / 5 |
Total | 16 / 20 Passed |