Content
86%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A well-structured, token-efficient triage workflow with a concrete failure taxonomy and sensible guardrails against overstating flakes as regressions. The main gap is actionability detail: harness detection and focused-rerun steps describe what to do without giving runnable command forms, and there is no explicit fix-verification step.
Suggestions
Add runnable example commands for the rerun and scope-narrowing steps, e.g. `xcodebuild test -only-testing:TestTargetTests/TestClass/testMethod` and `swift test --filter TestClass.testMethod`.
Make harness detection concrete: detect the harness by the presence of `Package.swift` (SwiftPM) vs. `.xcodeproj`/`.xcworkspace` (Xcode) instead of only naming the two commands.
Close the workflow loop with an explicit verification checkpoint after a fix, e.g. rerun the originally failing test and confirm it passes before reporting the regression as resolved.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is lean and assumes Claude's competence: no concept explanations, no tool tutorials, just directives like "Use `xcodebuild test` for Xcode-based projects" and "Distinguish compilation failures from test execution failures". Every section earns its tokens, matching the top anchor rather than the 4 anchor ("minor instances of over-explanation"). | 5 / 5 |
Actionability | Concrete commands are named (`xcodebuild test`, `swift test`) and the six-category failure taxonomy plus output checklist are specific, but the rerun and scope-narrowing steps lack executable specifics (no example flags like `-only-testing:` or `--filter`, no example commands). This sits between the 3 anchor (missing key details) and the 5 anchor (copy-paste ready commands) but noticeably above the midpoint. | 4 / 5 |
Workflow Clarity | A clear five-step sequence (detect harness → narrow scope → classify → rerun → summarize) with a built-in decision loop ("Use focused reruns when a specific case fails... without new information") and an output checklist. It misses a 5 because there is no explicit verify-the-fix checkpoint; it exceeds 3 because the sequence and checkpoints are explicit rather than implicit, and no destructive/batch validation cap applies. | 4 / 5 |
Progressive Disclosure | The skill is under 50 lines, needs no external references (none exist in the bundle), and is organized into four clearly labeled sections (Quick Start, Workflow, Guardrails, Output Expectations) with nothing inlined that belongs in a separate file — matching the simple-skill exception for a top score. | 5 / 5 |
Total | 18 / 20 Passed |