Content
81%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A highly actionable, well-sequenced evaluation skill: every rule check is a verbatim command, the gate logic is unambiguous, and validation/feedback loops are explicit. Its weakness is token efficiency — the same justifications (verbatim-command notes, alias deprecation, 89% cap) are repeated multiple times, and the rules table could be split into a reference file.
Suggestions
State the no-shell-pipes/verbatim justification for the awk commands once (e.g., in a note above the rule table) instead of repeating it in rules #5, #6, #7, #9, and #10.
Consolidate the deprecated `bug-analysis` alias handling into a single paragraph — it is currently explained in the metadata tag comment, the mode table, a dedicated alias-handling paragraph, and again in the Fix Mode section.
Move the 11 deterministic rule-check commands to a one-level-deep reference file (e.g. references/plan-rules.md), keeping a rule-number/pass-fail summary table in SKILL.md to cut context cost for code and analysis invocations that never run them.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Mostly efficient, load-bearing tables, but with real repetition: the justification "Single awk command with no shell pipes, so the table row executes verbatim" is repeated across five rules, the deprecated `bug-analysis` alias is explained in four separate places (metadata tag comment, mode table, alias-handling paragraph, Fix Mode section), and the 89% cap is restated in the callout, rule intro, Step 3, and Output Format. This fits anchor 3 (could be tightened) better than anchor 4's "minor instances". | 3 / 5 |
Actionability | Fully executable guidance: copy-paste-ready `test -f`, `grep`, and single-command `awk` one-liners for all 11 deterministic rule checks, an exact output-format template with a worked table skeleton, explicit dimension weights, and numeric score thresholds. Specific examples cover the plan/code/analysis common cases — anchor 5. | 5 / 5 |
Workflow Clarity | Plan mode is explicitly sequenced (Step 1 LLM dimensional scoring → Step 2 deterministic rule checks → Step 3 combined gate) with a hard validation checkpoint ("A failing rule is a blocker the gate must surface even if the LLM dimensional score is high"), an iteration protocol that re-runs the assessment after each of up to 2 iterations, and a Post-Fix Re-Assessment feedback loop. This matches anchor 5's explicit validation, feedback loops, and checklists. | 5 / 5 |
Progressive Disclosure | No bundle files exist, and the single-file body is well navigated via a Contents list and clean section headers; the Output Format and Fix Mode sections are cleanly gated ("omit for code/analysis", "Skip this section entirely if not in Fix Mode"). The ~60-line rule-check table with long inline awk commands is a plausible candidate for a one-level-deep reference file, which keeps this at anchor 4 rather than 5. | 4 / 5 |
Total | 17 / 20 Passed |