Content
80%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The body is lean, well-structured, and highly actionable, with executable gh commands and a concrete rating rubric. The main weakness is the Mode B batch path: it lacks a post-apply verification and error-recovery loop, which caps workflow clarity at 3 per the rubric's batch-operation rule, and the amended-rating path (removing the old label) is never shown. A duplicated repo name in the labels section is also a small factual defect.
Suggestions
Add a verification loop to Mode B: after the batch pass, re-run the `gh issue list` jq filter and confirm it returns an empty list, and retry or report the issue numbers whose `gh issue edit` failed — this satisfies the feedback-loop requirement for batch operations.
Show how an amended rating is corrected: `gh issue edit --remove-label "complexity:3"` before adding the new value, since Mode A explicitly offers neighbouring values but `--add-label` alone would leave stale labels on the issue.
Fix 'Six, on both englishstreetventures/englishstreetventures.com and englishstreetventures/englishstreetventures.com' — the same repo is named twice where two different repos were presumably intended, which makes the label scope ambiguous.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Lean and disciplined throughout — no explanation of concepts Claude already knows, and every section (timing rationale, rubric table, two modes, labels, consumers) carries system-specific constraints that could not be inferred, e.g. "the rating happens at issue time, never at pull-request time" and "Never let the agent that did the work rate the work". Nothing could be trimmed without losing real information, so this fits the 5 anchor rather than 4's 'minor instances that could be trimmed'. | 5 / 5 |
Actionability | Concrete, copy-paste-ready commands: `gh issue edit 412 --repo ... --add-label "complexity:3"` and the full `gh issue list --json number,title,body,labels --jq ...` filter, plus a concrete Fibonacci rubric table with example shapes. Not 5 because of minor gaps: amending a rating would require `--remove-label`, which is never shown, and the labels section says "on both englishstreetventures/englishstreetventures.com and englishstreetventures/englishstreetventures.com" — the same repo named twice where two were presumably intended. Well above 3, since the shown commands are fully executable, not pseudocode. | 4 / 5 |
Workflow Clarity | Mode A is clearly sequenced with an explicit confirm/amend checkpoint and an unattended fallback, but Mode B is a batch operation — "Do them in one pass over a listing" of up to 200 issues — with only a pre-selection filter and no post-apply verification or error-recovery loop (no re-check that the listing is now empty, no handling for failed edits). Per the rubric's explicit cap, a batch workflow without validation cannot score above 3, which takes precedence over the otherwise clear sequencing; anchor 4's 'most checkpoints present' is not reached on the batch path. | 3 / 5 |
Progressive Disclosure | Self-contained with well-organized sections (Why the timing, The rubric, Mode A, Mode B, The labels, What consumes this) and no bundle files; nothing that belongs in a separate file is inlined, and the only external mentions (`@tools/pr-metrics`, `wiki/conventions/session-metrics.md`) are clearly signaled one-level-deep informational pointers. Fits the 'well-organized sections' exception for self-contained skills; not 4 because no content is misplaced. | 5 / 5 |
Total | 17 / 20 Passed |