Content
83%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A tight, actionable operating manual for a measurement-gated improvement loop, with strong workflow sequencing and explicit validation/guardrails appropriate to a destructive-capable campaign. Conciseness and progressive disclosure are very good but not perfect, with some inlined detail and minor restated rationale.
Suggestions
Move the exact agent-mode runner command (and the exact slice-selection rule) into the body so both probe and agent modes are equally copy-pasteable.
Consider splitting the detailed journal/PR-evidence format fully into references/campaign-journal.md and linking to it rather than restating the 'before/after scores in the body' format inline in step 5.
Trim restated guardrail rationale (e.g. 'that's why the no-regression sample is mandatory', 'changing the exam and the answer together proves nothing') since Claude can infer the why from the rule.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Mostly lean and assumes Claude's competence — it skips explaining what MCP or eval harnesses are and uses tight command strings — but the guardrails and failure-modes sections include some defensive restating of rationale ('Anything else... → stop', 'that's why the no-regression sample is mandatory') that could be trimmed. | 4 / 5 |
Actionability | Provides concrete, copy-pasteable commands (probe.ts invocation, the pnpm dev:hono local recipe with env vars, named analytics tools) and a precise allowlist; minor gaps are the absense of the exact agent-mode command and the benchmark slice selection being left implicit. | 4 / 5 |
Workflow Clarity | The 'One iteration' section is a clearly sequenced 6-step procedure with explicit validation (step 4: re-run slice + no-regression sample, keep only if metric improves and nothing degrades) and feedback loops (discard → journal → move on), plus hard guardrails and a kill switch — matching the top anchor for batch/destructive workflows. | 5 / 5 |
Progressive Disclosure | SKILL.md is a well-structured overview with one clearly signaled one-level-deep reference (references/campaign-journal.md, which exists and holds the journal/PR-evidence detail); the cross-link to the intent-clusters skill is also clearly signaled. Not a 5 because some reference-worthy detail (failure modes, the full journal format) is inlined rather than split out. | 4 / 5 |
Total | 17 / 20 Passed |