CtrlK
BlogDocsLog inGet started
Tessl Logo

ab-test-setup

Use when designing an A/B or split test: define the hypothesis, control and variants, estimate sample size, verify tracking, and predeclare metrics and stopping rules.

70

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-engineered procedural skill: gated workflow with explicit validation checkpoints, feedback loops, refusal conditions, and executable statistical code. Its weaknesses are minor — some duplicated gating/principles content that could be tightened, and a fully inline structure with no offloading of illustrative material to reference files.

Suggestions

Collapse the 'Key Principles (Non-Negotiable)' section or merge it into the existing gates: its items (one hypothesis, one primary metric, commit before launch, no peeking) are already stated in sections 3, 6, 7, and 8, so the repetition costs tokens without adding guidance.

Give the remaining un-executed steps concrete form: for the sample-ratio-mismatch check, name a specific test (e.g., a chi-square goodness-of-fit test on assignment counts) the way the sample-size calculation is made concrete with runnable code.

Consider moving the sample-size calculation example and worked example into a references/ file (e.g., references/examples.md) linked from the main body, keeping SKILL.md as a tighter overview of the gated procedure.

DimensionReasoningScore

Conciseness

The body is almost entirely lean imperative bullets with no explanations of concepts Claude already knows, but there is trimmable redundancy: the "Key Principles (Non-Negotiable)" section ("One hypothesis per test", "One primary metric", "No peeking") restates rules already established in sections 3, 6, and 8, and the gating language repeats across sections 3, 7, and 8.

4 / 5

Actionability

Concrete, executable guidance dominates: a complete copy-paste Python sample-size calculation with expected output ("14745 observations per variant"), a worked example with concrete units and guardrails, and specific verification steps ("compare a sample of 5+ events per variant", "stable event/transaction ID"). Minor gaps remain — e.g., guardrail dashboard/alert setup and the sample-ratio-mismatch check are named but not given executable form.

4 / 5

Workflow Clarity

The process is explicitly sequenced (numbered sections 1-8) with hard gates ("Hypothesis Lock (Hard Gate)", "Execution Readiness Gate (Hard Stop)"), a pre-gate tracking verification checklist, error-recovery loops ("If any of the above fails, stop and resolve it before Gate 8"; "If any item is missing, stop and resolve it"), and a refusal-conditions section — matching the anchor for clear sequence with explicit validation, feedback loops, and checklists.

5 / 5

Progressive Disclosure

No bundle files exist (references/, scripts/, assets/ are all absent), so all content is inline in a single ~280-line file. Internal structure is strong — numbered sections, clear headers, a decision table — and nothing is deeply nested or buried, but a few blocks (the sample-size code example, worked example, limitations) could arguably live in reference files to keep the main body closer to an overview, which keeps it below the top anchor.

4 / 5

Total

17

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that pairs an explicit trigger clause with a specific, comprehensive list of capabilities in third person. Its only gap is modest synonym coverage for the trigger terms.

DimensionReasoningScore

Specificity

The description enumerates five concrete, distinct actions — "define the hypothesis, control and variants, estimate sample size, verify tracking, and predeclare metrics and stopping rules" — giving comprehensive coverage of the skill's capability surface rather than vague domain language.

5 / 5

Completeness

It explicitly answers both questions: what the skill does ("define the hypothesis, control and variants, estimate sample size, verify tracking, and predeclare metrics and stopping rules") and when to use it via the concrete "Use when designing an A/B or split test" trigger clause.

5 / 5

Trigger Term Quality

"A/B or split test" are the natural phrases users would say, supported by accessible terms like hypothesis, sample size, metrics, and stopping rules. Common synonyms such as "experiment", "A/B testing", or "conversion" are absent, so it falls just below the comprehensive anchor.

4 / 5

Distinctiveness Conflict Risk

The A/B-testing niche is clearly staked out with domain-specific triggers (hypothesis, MDE-adjacent vocabulary, stopping rules, tracking verification) that would not naturally fire for analytics, feature-flag, or general statistics skills, keeping conflict risk minimal.

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
boisenoise/skills-collections
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.