Content
78%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A token-efficient, well-sequenced experimental-design checklist that covers setup, readout, and anti-patterns without padding. Its main weakness is actionability: the sample-size/duration calculation and the required versioned JSON output are mandated but never specified, leaving Claude to fill in method details.
Suggestions
Add the sample-size/duration calculation method (e.g., baseline rate, MDE, power/alpha defaults, or a formula/reference) so step 3 is executable rather than just mandated.
Include a minimal example of the versioned JSON setup/readout schema referenced in step 7 so the required output shape is concrete.
Name the specific platform experiment tools to check in step 2 (e.g., Google Ads campaign drafts/experiments, Meta A/B test tool) instead of the generic "platform experiment tools".
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The 20-line body is lean and assumes Claude's competence — it names required design elements and pitfalls ("Do not repeatedly peek and stop on a favorable result, call underpowered noise a winner") without explaining any concepts Claude already knows, so every token earns its place. | 5 / 5 |
Actionability | The guidance names concrete artifacts ("treatment, control, randomization unit, population, primary metric, guardrails, minimum effect, and stopping rule") but leaves key execution details missing: no method or formula for "Calculate sample and duration", no example of the required "versioned JSON" output format, and no named platform experiment tools. This fits the anchor for some concrete guidance but incomplete with missing key details, rather than mostly-executable guidance. | 3 / 5 |
Workflow Clarity | A clear 7-step sequence with an explicit validation gate — "verify assignment integrity and data completeness before estimating effect and uncertainty" — plus pre-registration of analysis and thresholds. It stops short of the top anchor because there is no feedback loop describing what to do when the verification checks fail. | 4 / 5 |
Progressive Disclosure | The skill is under 50 lines with no bundle files and no inlined content that belongs in separate references; a short, well-organized single-purpose body fully satisfies the simple-skill exception for progressive disclosure. | 5 / 5 |
Total | 17 / 20 Passed |