Content
96%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
An exemplary instruction-only skill: dense with project-specific knowledge, copy-paste-ready commands, explicit validation gates, and a memorable catalog of past failure modes that doubles as an error-check checklist. The only structural note is that progressive disclosure relies entirely on external repo docs rather than bundle files, which is appropriate here but leaves minor navigation gaps.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Every section carries project-specific knowledge Claude could not have (the two past wrong findings, `explore_share` null semantics, median 3.6M vs mean 15.4M skew, the `at-open` container death) with zero general-concept padding — e.g. "Nothing else counts as an analysis — a table restated in prose is not one." Not 4 because there is no passage of over-explanation that could be trimmed without losing operational content; the median-vs-mean and correlation discussion is grounded in this corpus's numbers, not textbook explanation. | 5 / 5 |
Actionability | Fully executable commands are given — `bun run --cwd tools/pr-metrics report -- --coverage` / `--waste` / `--json`, and the copy-paste `gh pr diff 882 ... | grep -E '^[+-]\s*"...'` — plus concrete decision tables (declared complexity vs spend) and an exact output spec ("the number, the sample size it rests on, and what would have to be true for it to be wrong"). Not 4 because the commands and examples cover the common cases with no missing key details. | 5 / 5 |
Workflow Clarity | Explicit gated sequence — Step 0 coverage ("write them down before continuing", "If nothing is rated, say so at the top of your answer"), Step 1 run, the four traps as an interpretation checklist, Step 2 report — with per-claim validation built in (each claim must name what would overturn it, and exclusions must be reported). Not 4 because validation checkpoints are explicit and complete, including error-recovery guidance for the exact failure modes that previously occurred; this is a read-only analysis skill so the destructive-operation cap does not apply. | 5 / 5 |
Progressive Disclosure | Well-organized single-file structure with clear section headers and one-level-deep, clearly signaled external pointers ("`wiki/conventions/session-metrics.md` holds the schema... `tools/pr-metrics/README.md` has the commands"); no bundle files exist, so everything operational is appropriately inline. Not 5 because at ~120 lines the body exceeds the under-50-line simple-skill case, and some detail (e.g. the per-metric schema semantics) is deferred but could be more prominently navigable; not 3 because no content that clearly belongs in a separate file is inlined and references are not buried. | 4 / 5 |
Total | 19 / 20 Passed |