Content
78%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A strong, expert instruction-only skill: highly actionable specific guidance and excellent progressive disclosure with seven real, well-signaled reference files. The main weakness is conciseness — recurring philosophical framing and re-explanation of statistical concepts Claude already knows add tokens that could be trimmed.
Suggestions
Tighten the intro and 'Closing: when in doubt' sections: drop the rhetorical framing about the default state of experimentation and keep only the decision-relevant guidance, cutting roughly 15-20% of tokens.
Compress the re-explanations of known statistical concepts (multiple-comparisons false-positive math, peeking inflation percentages, novelty/primacy definitions) into one-line reminders that point to the decision rule rather than re-deriving them.
Promote the pre-experiment readiness and post-experiment decision steps into a single explicit numbered workflow with validation checkpoints in SKILL.md, so the lifecycle sequence is unambiguous rather than implied across sections.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is dense with genuinely expert, specific guidance, but it also re-explains concepts Claude already knows (the multiple-commissions math, peeking false-positive inflation, novelty/primacy effects) and includes philosophical framing ('The default state of experimentation in most companies is sloppy', the 'Closing: when in doubt' section) that could be trimmed without losing actionability. | 3 / 5 |
Actionability | Highly concrete and specific for an instruction-only skill: 'Pick exactly one primary metric. Pick three to five guardrails', 'Two weeks is the conventional minimum for any UI/UX experiment', 'Maximum duration. Usually four to six weeks', and the vendor question 'What is your variance estimator for ratio metrics?' — absence of code is not penalized because the guidance is directly actionable. | 5 / 5 |
Workflow Clarity | The 12-consideration framework plus the lifecycle ordering (readiness -> hypothesis -> sizing -> duration -> running -> interpretation -> decision) and ranked inconclusive-resolution paths give a clear sequence with checkpoints (falsifiability test, pre-commitment, referenced readiness checklist), though it reads more as a playbook than a strictly validated step sequence. | 4 / 5 |
Progressive Disclosure | Well-structured overview with one-level-deep references to seven verified bundle files (hypothesis-templates, sample-size-tables, common-failures, results-interpretation-checklist, platform-comparison, pre-experiment-readiness-checklist, post-experiment-decision-framework), each clearly signaled via markdown links and a dedicated 'Reference files' section with descriptions. | 5 / 5 |
Total | 17 / 20 Passed |