Content
81%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The body is highly actionable and operationally safe: executable commands throughout, a single-script ownership model with explicit blocked-state handling, and a report contract that removes all ambiguity from the inverted pass/fail semantics. Its weaknesses are duplication — the modes, inverted semantics, and output layout are each stated three to four times — and a monolithic 265-line structure where troubleshooting and parameter reference material would be better split into a one-level-deep reference file. Both issues are about token efficiency, not correctness.
Suggestions
Collapse the mode duplication: keep the "Mode 1"/"Mode 2" sections with their commands and delete the redundant restatements in "Requirements" and "What It Does", which repeat the same fail-without-fix / pass-with-fix logic almost verbatim.
State the inverted pass/fail semantics once (the table plus one warning sentence); remove the repeated reminders in Step 3, Step 4, and the intro, trusting the initial statement to carry.
Move the Output Files directory tree, Troubleshooting table, and Optional Parameters into a one-level-deep reference file (e.g. references/parameters.md) and link to it, shrinking SKILL.md to the workflow core.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is mostly project-specific knowledge Claude could not know (runner scripts, paths, parameters, output contracts), which earns its tokens — but the same facts are repeated several times: the two modes are explained in "Workflow Step 1", again in "Mode 1"/"Mode 2", again in "Requirements", and again in "What It Does"; the inverted pass/fail semantics appear four times (the table, "NEVER say 'verification passed'", the Step 3 reminder, and the Step 4 bullet list); and the output-file layout is described twice ("Output Files" table and the example tree). This matches anchor 3 — mostly efficient with unnecessary duplication that could be tightened — rather than 2, since almost no space is spent on concepts Claude already knows. | 3 / 5 |
Actionability | Every instruction is copy-paste executable: exact pwsh invocations with real parameter values ("pwsh .github/skills/verify-tests-fail-without-fix/scripts/verify-tests-fail.ps1 -Platform android -TestFilter 'Maui12345' -RequireFullVerification"), the referenced script actually exists in the bundle (scripts/verify-tests-fail.ps1), the expected output markers are shown verbatim, and the Optional Parameters section gives concrete syntax for every flag including the frozen-fixture "$(git rev-parse HEAD^)" variant. Not below 5: the common cases (auto-detect, explicit filter, full verification) are all covered with working commands. | 5 / 5 |
Workflow Clarity | The four-step workflow is clearly sequenced with strong validation checkpoints: mode determination from the git diff, a single script owning all fix-file reversion (with an explicit prohibition on manual git cleanup), unambiguous terminal markers (VERIFICATION PASSED / FAILED / error-timeout → Blocked), explicit background-session handling, and a report contract that enumerates what each marker means in each mode. The operation is destructive (reverting fix files), but the guardrails — the script owning transitions, reporting Blocked rather than cleaning up, and the troubleshooting table mapping failure → cause → solution — provide exactly the feedback loops the top anchor requires. | 5 / 5 |
Progressive Disclosure | The body is well-sectioned with clear headers (Supported Test Types, Workflow, Expected Output, Output Files, Troubleshooting, Optional Parameters) and its one script reference is real and matches the bundle layout (scripts/verify-tests-fail.ps1). It does not reach 5 because the file is a ~265-line monolith: content that would sit naturally in a one-level-deep reference — the output-file directory tree, the troubleshooting table, and the full parameter reference — is inlined in SKILL.md, and no reference files exist in the bundle. It is comfortably above 3 since what is present is organized and navigable. | 4 / 5 |
Total | 17 / 20 Passed |