Content
70%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
This is an exceptionally actionable and rigorously sequenced gate skill — concrete commands, exhaustive exit-code semantics, and validation at every phase — undermined by verbosity and zero progressive disclosure. A 750-line single file inlines per-engine material that belongs in separate reference files, and repeats both instructions and design-rationale commentary that add tokens without adding instruction.
Suggestions
Move the per-engine runner instructions (Phase 2's Godot, Unity, and Unreal sections, ~200 lines) into separate reference files such as references/godot-runner.md, references/unity-runner.md, references/unreal-runner.md, keeping only the engine-selection dispatch and verdict-relevant exit codes in SKILL.md.
Replace the six verbatim repetitions of 'For any selected item, ask the user to briefly describe what failed before generating the report' with a single rule stated once before the batch definitions.
Cut or compress the design-rationale blockquotes ('A bare halt is the wrong shape here…', the NOT ASSESSED ranking defense, 'This is a change… and it is deliberate') into one-line statements of the rule — the reasoning essays are commentary, not instruction, and cost significant tokens.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is noticeably verbose across several padded sections: the instruction 'For any selected item, ask the user to briefly describe what failed before generating the report' is repeated verbatim ~6 times across Batch 1/2/3 and the platform batches, and multi-line blockquote essays ('A bare halt is the wrong shape here…', the NOT ASSESSED ranking defense, 'This is a change in where unconfirmed NOT RUN lands, and it is deliberate') are design-rationale commentary rather than instructions. It is above level 1 because the engine-specific exit codes, timeouts, and gotchas are genuinely non-obvious knowledge Claude does not already have; it is below level 3 because the repetition and rationale-padding are substantial and easily trimmed. | 2 / 5 |
Actionability | Guidance is fully executable: exact copy-paste bash commands with timeout wrappers for all three engines and all three OSes, per-engine exit-code-to-verdict mappings (e.g. gdUnit4 0/101/100/105/103/104/1), concrete resolution rules for placeholders ('$(pwd -W 2>/dev/null || pwd) gives the C:/… form in Git Bash'), explicit AskUserQuestion batch definitions, and a complete report template. The common cases (Godot, Unity EditMode+PlayMode, Unreal) are each covered with runnable commands, matching the top anchor; it is not level 4 because no significant execution gap was found. | 5 / 5 |
Workflow Clarity | Six phases are clearly sequenced (detect setup → run tests → coverage → manual checks → report → gate) with explicit validation checkpoints at every phase: exit-code interpretation rules, 'a timeout (exit 124) is a gate FAILURE, never a pass', deletion of stale results files before runs, a first-matching-rule-wins verdict table, and write-only-after-approval gating. Error-recovery feedback loops ('fix the failures and run /smoke-check again') are explicit, matching the top anchor; it is not level 4 because checkpoints are present with no material gaps. | 5 / 5 |
Progressive Disclosure | The file has clear, navigable phase headers, so structure is not minimal — but ~200 lines of per-engine runner instructions (Phase 2's Godot/Unity/Unreal sections) and the long report template are inlined in a 750-line monolithic SKILL.md with no references/ bundle at all, matching the 'some structure but… content that should be separate is inline' anchor. It is above level 2 (headers make it navigable, and external doc references are named, e.g. docs/engine-reference/unreal/current-best-practices.md) but below level 4 because content that clearly belongs in per-engine reference files sits inline. | 3 / 5 |
Total | 15 / 20 Passed |