Content
85%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A strong, highly actionable instruction skill: the tiering rubric is concrete and executable (real case-matching code, explicit ordering, exact output token), and the three-step workflow is well-gated with evidence citation and regression-tested evals. The main weaknesses are moderate repetition of the severity-vs-confidence/policy-free framing and, most importantly, progressive disclosure: the body leans on multiple bundle and repo paths that are not actually present alongside the SKILL.md, leaving its references unresolvable and its consumer-integration detail inlined.
Suggestions
Ship the referenced files (at minimum `scripts/eval/golden/severity-tiering.jsonl` and the eval harness, plus the consumer-integration docs like conventional-comments.md / review-config.md) inside the skill's references/ or scripts/ directories, or rewrite those mentions as self-contained descriptions so no reference dangles.
Move the repo-specific consumer material (the 'Mapping to a reviewer's blocking flag' crosswalk and 'How Callers Consume the Tier' sections) into a reference file linked one level deep, keeping SKILL.md as the rubric + output format core.
State the severity-vs-confidence orthogonality and the policy-free stance once (the opening blockquote) and drop the re-statements in 'How Callers Consume the Tier' to tighten token efficiency.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The core sections (tier table, axes, Step 1–3, output format) are dense and earn their tokens — executable case blocks, tables, and an exact format template with no explaining of concepts Claude already knows. It is not level 5 because points repeat: the severity-vs-confidence distinction appears in the opening blockquote and again in "How Callers Consume the Tier" ("it is advisory, like confidence, and it is policy-free"), and the policy-free point is made twice. Not level 3 because the padding is minor relative to genuinely needed rubric content. | 4 / 5 |
Actionability | Fully executable guidance: copy-paste-ready `case "$path" in ... esac` blocks for the exclusion gate and the path floor, an explicit combination rule (`severity = max(base_tier, path_floor, escalator_minimums)` over `low < medium < high < critical`), specific escalator shapes with exact minimums, and a byte-exact machine-readable output format (`## Severity: <critical|high|medium|low>`). It is not level 4 because no key execution detail is missing — the common cases are covered by the tier table plus escalator table plus the tricky-case golden set. | 5 / 5 |
Workflow Clarity | The multi-step process is clearly sequenced (Step 1 axes with reachability cap → Step 2 with the exclusion gate explicitly run first, then floor, then escalators → Step 3 max-combination), with validation checkpoints in the output format (each escalator must cite "file:line", exclusion and floor matches shown as ✓/✗) and a feedback loop via the regression-tested golden eval suite that locks tricky cases (test-path destructive statement, billing path floor, dead-path catastrophic bug). The destructive/batch cap does not apply — the skill is advisory and performs no destructive or batch operations. Not level 4: checkpoints are explicit, not implicit. | 5 / 5 |
Progressive Disclosure | Internal structure is good (a Contents TOC with anchors, and the rubric section explicitly flagged "self-contained"), but the body references files that do not exist in the skill bundle — `scripts/eval/golden/severity-tiering.jsonl`, `scripts/eval/l2.mjs`, `scripts/record-comment-relevance.mjs`, `agents/shared/rules/conventional-comments.md`, `rubric-composition.md`, `review-config.md` — and no references/, scripts/, or assets/ directories are present, so the references are unresolvable. Additionally, the consumer-integration material (the blocking-flag crosswalk and caller-consumption sections) is inlined where it could live in a reference file. This matches the level-3 anchor (some structure, references present but not clearly resolvable, content that should be separate is inline); not level 4 because dangling paths are more than a minor organization gap. | 3 / 5 |
Total | 17 / 20 Passed |