Content
81%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A well-engineered procedural skill: gated workflow with explicit validation checkpoints, feedback loops, refusal conditions, and executable statistical code. Its weaknesses are minor — some duplicated gating/principles content that could be tightened, and a fully inline structure with no offloading of illustrative material to reference files.
Suggestions
Collapse the 'Key Principles (Non-Negotiable)' section or merge it into the existing gates: its items (one hypothesis, one primary metric, commit before launch, no peeking) are already stated in sections 3, 6, 7, and 8, so the repetition costs tokens without adding guidance.
Give the remaining un-executed steps concrete form: for the sample-ratio-mismatch check, name a specific test (e.g., a chi-square goodness-of-fit test on assignment counts) the way the sample-size calculation is made concrete with runnable code.
Consider moving the sample-size calculation example and worked example into a references/ file (e.g., references/examples.md) linked from the main body, keeping SKILL.md as a tighter overview of the gated procedure.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is almost entirely lean imperative bullets with no explanations of concepts Claude already knows, but there is trimmable redundancy: the "Key Principles (Non-Negotiable)" section ("One hypothesis per test", "One primary metric", "No peeking") restates rules already established in sections 3, 6, and 8, and the gating language repeats across sections 3, 7, and 8. | 4 / 5 |
Actionability | Concrete, executable guidance dominates: a complete copy-paste Python sample-size calculation with expected output ("14745 observations per variant"), a worked example with concrete units and guardrails, and specific verification steps ("compare a sample of 5+ events per variant", "stable event/transaction ID"). Minor gaps remain — e.g., guardrail dashboard/alert setup and the sample-ratio-mismatch check are named but not given executable form. | 4 / 5 |
Workflow Clarity | The process is explicitly sequenced (numbered sections 1-8) with hard gates ("Hypothesis Lock (Hard Gate)", "Execution Readiness Gate (Hard Stop)"), a pre-gate tracking verification checklist, error-recovery loops ("If any of the above fails, stop and resolve it before Gate 8"; "If any item is missing, stop and resolve it"), and a refusal-conditions section — matching the anchor for clear sequence with explicit validation, feedback loops, and checklists. | 5 / 5 |
Progressive Disclosure | No bundle files exist (references/, scripts/, assets/ are all absent), so all content is inline in a single ~280-line file. Internal structure is strong — numbered sections, clear headers, a decision table — and nothing is deeply nested or buried, but a few blocks (the sample-size code example, worked example, limitations) could arguably live in reference files to keep the main body closer to an overview, which keeps it below the top anchor. | 4 / 5 |
Total | 17 / 20 Passed |