Content
78%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
Highly actionable and information-dense, with executable commands, precise output contracts, and good error-handling guidance. Its main weakness is structure: the SKILL.md is a long monolithic reference that would benefit from splitting advanced flag deep-dives into separate reference files rather than inlining everything.
Suggestions
Move the deep-dive material for advanced modes — 'Shortlist with rank uncertainty' (--eb, rank diagnostics), '--touching' gating internals, and 'Owner evidence' — into one-level-deep reference files (e.g. references/shortlist.md, references/touching.md), leaving SKILL.md a concise overview with pointers.
Break the dense single-paragraph sections into tables or bullet lists (e.g. the '--touching' and '--priors' sections) so key facts like field names and matcher semantics are scannable rather than buried mid-paragraph.
Add one worked end-to-end example for the shortlist flow (command, expected stderr rank_report line, and how to act on pTopK/poth), since that section currently describes fields without showing a concrete invocation.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is information-dense and assumes Claude's competence — no generic concepts are explained, and nearly every sentence carries spn-specific contract details. Minor trimming is possible: sections like '--touching' and 'ownerEvidence' are long clause-stacked single paragraphs that could be tightened or tabulated, which keeps this below a 5. | 4 / 5 |
Actionability | Copy-paste-ready commands throughout: the core loop ('spn forks list owner/repo | jq -c ...'), a concrete repo-search pipeline, exact flag syntax for every mode, an exit-code decision table, and field-by-field tables for shortlist and owner-evidence records. The common cases are covered with executable examples. | 5 / 5 |
Workflow Clarity | The main flows are clearly sequenced (first check `command -v spn`, fall back to `gh api` if absent; stream and filter with jq) with error-recovery checkpoints — exit 2 'do not retry' vs. retryable errors, rate-limit `retry_after_seconds`, 'check the run output for `cluster_skip` warnings', and 'Check `ebRegime` in the `rank_report` first'. Minor gaps remain in sequencing the advanced shortlist/rank-diagnostics flow, so it sits just below a 5. | 4 / 5 |
Progressive Disclosure | Section headers, tables, and a Quick Reference make navigation decent, but there are no bundle files at all: reference-grade detail — the `--shortlist` rank-uncertainty math and empirical-Bayes internals, the `--touching` gating/batch-compare internals, and the `ownerEvidence` schema — is inlined in SKILL.md where it belongs in one-level-deep reference files. This matches 'content that should be separate is inline' better than a 4. | 3 / 5 |
Total | 16 / 20 Passed |