Content
92%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A well-engineered orchestrator skill: fully executable commands with explicit verdict contracts at every step, a clearly gated phase workflow with checkpointing and error-recovery, and exemplary progressive disclosure with nine real one-level-deep references. The only weakness is minor redundancy where a few rules and inline mechanisms are restated or belong in the references, keeping conciseness just below the top anchor.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is dense rules and verbatim commands with no explanations of concepts Claude already knows, but a few rules are restated across sections (pattern-selection appears in Non-negotiable rules, step 7, and the investigate section; the "show the evidence block" rule appears twice) and some inline mechanisms (the merged-fix denominator logic, Last-seen rendering like `5d ago (Jul 24)`) exceed the file's own "invariant here, mechanism in the reference" split. These are minor trimmable instances, matching level 4 rather than the every-token-earns-its-place level 5. | 4 / 5 |
Actionability | Every stage carries a copy-paste-ready command with flags and expected output shapes: `resolve-test-key.js '<whatever they gave>'`, `triage-history.js --test-key ... --lookback-days 14`, `checkpoint.js --triage-id <id> --init --test-key '<key>'`, `record-diagnosis.js --pr <n> --outcome fix-test`. Verdict handling (`stop: true`, `cause: "missing-api-key"`, `no-results`, `open-attempt-in-flight`) is specific, and the exact table format to present is shown. Specific examples cover the common cases. | 5 / 5 |
Workflow Clarity | The multi-step process is explicitly sequenced (resolve identity → history → checkpoint init → prior-triage check → pattern table → selection → evidence → diagnosis → fix approach gate → reproduce → record outcome) with validation throughout: API-key pre-flight, `--read` validation on resume, the checkpoint gate that refuses `phase=done` without a recorded diagnosis, the RED bar for regression tests, and error-recovery paths (candidates → ask which, 403/expired report handling, script-fallbacks reference). | 5 / 5 |
Progressive Disclosure | SKILL.md is a genuine overview holding invariants while nine real reference files (all verified to exist) each own a complete mechanism, every one clearly signaled with what it contains ("owns the verdict table, the run-it offer for no-results, what local evidence cannot answer"). References are one level deep with only lateral sibling links, the scripts table points to references/scripts.md for flags and output contracts, and the file states its own split policy ("this file holds the invariant, the reference holds the mechanism"). | 5 / 5 |
Total | 19 / 20 Passed |