Content
77%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The body is an exceptionally actionable, well-sequenced runbook with genuine validation and feedback loops — every command is executable and every judgment call is named. Its weaknesses are structural: the read-only and retry-disclosure rules are repeated many times, and the large inline report template and coverage table belong in reference files, leaving SKILL.md as a monolith.
Suggestions
Move the ~80-line REPORT.md template and the Phase 4c coverage table into a references/ file (e.g. references/report-template.md and references/coverage-matrix.md), keeping only the section outline and a worked row or two in SKILL.md.
State the read-only rule once in Global rules (bolded) and the retry-disclosure rule once in Phase 4a, then reference them from the checklist and report template instead of restating them verbatim.
Trim the example report template to a few representative rows (e.g. one all-pass, one skip, one retry) rather than enumerating all 15 coverage areas inline.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Mostly efficient — the operational detail (pytest flags, env var sourcing, whitelist semantics) is genuinely project-specific — but there is notable repetition: the read-only rule is stated at least five times ("**This skill is read-only. It does not change any code... ever**", "**Read-only. Make no edits.**", "This skill only flags them — it does not edit prod code", "the skill never made any edits", checklist "No files edited outside..."), and the retry-disclosure rule is stated three times in Phase 4a plus again in the checklist and report template. The ~80-line inline report template could be tightened. This matches 'Mostly efficient but includes some unnecessary explanation or could be tightened'; it is not 4 because the duplication is systematic rather than a minor instance. | 3 / 5 |
Actionability | The guidance is fully executable and copy-paste ready: exact bash commands ("uv run ./checks.sh --agent-mode 2>&1 | tee \"${OUT}/checks.log\""), the exact pytest invocation with a justified "-o \"addopts=\"" override, a complete runnable Python status-check script, concrete grep pipelines, and a decision tree keyed on real error signatures ("model not found", "404", "model_not_found"). This matches the anchor 'Fully executable; copy-paste ready code or commands; specific examples cover the common cases'. | 5 / 5 |
Workflow Clarity | Six clearly sequenced phases with explicit validation checkpoints and feedback loops: exit codes are captured with a "don't bail — keep going" rule, transient failures are retried once in isolation with a reclassification rule ("a test that passes on retry is transient/flaky... a test that fails again is reproducible"), ambiguous causes get a mandatory escalation rule, and a final checklist verifies every phase. This matches the anchor 'Clear sequence with explicit validation steps; feedback loops for error recovery; checklists for complex processes'. The destructive/batch cap does not apply since the skill is read-only and validation is pervasive. | 5 / 5 |
Progressive Disclosure | No bundle files exist (no references/, scripts/, or assets/ directories), so all ~430 lines live inline in SKILL.md. Headers give the document decent structure, but content that clearly belongs in a separate file is inlined — the ~80-line REPORT.md template and the long Phase 4c coverage table are classic reference-file material — and the one cross-reference ("specs/projects/code_tools/cross_os_spawn_checklist.md") points outside the bundle. This matches the anchor 'Some structure but could be better organized; content that should be separate is inline'; it is not 2 because the sectioning is real and navigation within the file is easy. | 3 / 5 |
Total | 16 / 20 Passed |