Content
82%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A well-structured, actionable CRO framework that gives Claude concrete templates, scoring rubrics, and a copy-paste output format. Main weaknesses are minor verbosity in the statistical foundations section and a heavier-than-overview SKILL.md with limited use of reference files.
Suggestions
Move the statistical foundations section (significance explanation, sample-size quick-reference table, multiple-testing note) into a references file, keeping only the decision-relevant rules inline to improve progressive disclosure.
Add an explicit pre-launch validation gate to the workflow (e.g., "QA variant renders and event tracking before starting the clock") to strengthen feedback loops.
Tighten the significance explanation to the CRO-specific misconception note rather than re-deriving what a p-value is.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Mostly lean for the depth covered — hypothesis template, ICE/PIE scoring, sample-size table, decision matrix, and output template all earn their tokens — with minor over-explanation of significance and Bonferroni that Claude already knows. Not a 5 because the statistical foundations section could be trimmed; not a 3 because padding is minor. | 4 / 5 |
Actionability | Provides copy-paste-ready, specific guidance for an instruction skill: a hypothesis structure with a worked example, 1-10 ICE/PIE scoring, a sample-size quick-reference table, a decision matrix, and a complete output-format markdown template. Not a 4 because the common cases are covered with no real gaps. | 5 / 5 |
Workflow Clarity | Clear 10-step workflow and 4-phase sequence (Audit → Hypothesis → Test design → Decide) with a pre-launch parameters checklist and a ship/kill/extend decision matrix. Not a 5 because the workflow reads as a sequence rather than explicit validate→fix→retry feedback loops; not a 3 because checkpoints are mostly present. | 4 / 5 |
Progressive Disclosure | Well-organized sections with one clearly-signaled, one-level-deep reference (references/hypothesis-library.md) verified to exist. Not a 5 because SKILL.md is heavier than an "overview" — the statistical foundations and sample-size table could be split out — and only one reference file exists; not a 3 because structure and signaling are good. | 4 / 5 |
Total | 17 / 20 Passed |