CtrlK
BlogDocsLog inGet started
Tessl Logo

ab-test-setup

When the user wants to plan, design, or implement an A/B test or experiment. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "conversion experiment," "statistical significance," or "test this." For tracking implementation, see analytics-tracking.

68

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

The risk profile of this skill

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, highly actionable skill body that leverages a real bundled calculator and template references with clear progressive disclosure. Its primary weakness is content redundancy across the principles, mistakes, and intake sections that inflates the token budget without adding capability.

Suggestions

Consolidate 'Common Mistakes' into 'Core Principles' and 'Running the Test' (it re-states stopping early, changing mid-test, and testing too many things already covered there) to remove roughly 30 lines of duplication.

Merge the 'Task-Specific Questions' intake into the 'Initial Assessment' section so intake guidance lives in one place rather than three.

Drop or shrink the inline sample-size quick-reference table since the fuller, correct tables already live in references/sample-size-guide.md, keeping only a pointer and the calculator command inline.

DimensionReasoningScore

Conciseness

The body is mostly efficient and avoids explaining basics Claude knows, but it carries real redundancy: 'Common Mistakes' restates 'Core Principles' and 'Running the Test' (stopping early, changing mid-test), the intake territory is covered in 'Initial Assessment', 'Task-Specific Questions', and 'Proactive Triggers', and the sample-size quick-reference table partially duplicates the fuller table in the reference file.

3 / 5

Actionability

Provides copy-paste-ready executable commands for the bundled calculator ('python3 scripts/sample_size_calculator.py --baseline 0.05 --mde 0.20 --json'), a concrete hypothesis template with strong/weak examples, checklists, and full referenced templates (test plan, results doc, prioritization scorecard) — the common cases are covered.

5 / 5

Workflow Clarity

A clear end-to-end sequence (assess → hypothesize → size → design → run with pre-launch checklist → analyze with analysis checklist → document) with explicit pre-launch gates and 'don't peek' validation, but it lacks an explicit validate→fix→retry feedback loop for error recovery, which keeps it below a 5.

4 / 5

Progressive Disclosure

Clean overview body with well-signaled, one-level-deep references to real bundle files ('For detailed sample size tables...: See references/sample-size-guide.md', 'For templates: See references/test-templates.md') and executable logic split into scripts/sample_size_calculator.py; content is appropriately separated and easy to navigate.

5 / 5

Total

17

/

20

Passed

Description

86%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, well-structured description that clearly states what the skill does and provides excellent, comprehensive trigger coverage with explicit 'Use when' guidance. Its main weakness is generic action verbs in the 'what' clause and a few broad trigger terms that overlap with sibling marketing skills.

Suggestions

Replace the generic 'plan, design, or implement' verbs with the skill's concrete operations, e.g. 'Calculate sample size, write a hypothesis, design variants, and analyze statistical significance for A/B tests.'

Tighten broad triggers like 'test this', 'experiment', and 'hypothesis' with context (e.g. 'design an experiment', 'conversion hypothesis') to reduce overlap with page-cro and analytics-tracking.

DimensionReasoningScore

Specificity

Names the A/B-testing domain and three action verbs ('plan, design, or implement an A/B test or experiment'), but the actions are generic verbs rather than the distinct concrete operations (sample-size calculation, hypothesis authoring, results analysis) the skill actually performs, so it is not comprehensive enough for a 4.

3 / 5

Completeness

Explicitly answers both 'what' ('plan, design, or implement an A/B test or experiment') and 'when' with concrete trigger phrases, plus a disambiguation pointer ('For tracking implementation, see analytics-tracking').

5 / 5

Trigger Term Quality

Comprehensive coverage of natural user phrases including synonyms ('A/B test', 'split test'), variations ('test this change', 'test this'), and domain-natural terms ('multivariate test', 'statistical significance', 'hypothesis', 'conversion experiment'); file extensions are not applicable to this conceptual skill.

5 / 5

Distinctiveness Conflict Risk

Mostly distinct with an explicit disambiguation pointer to analytics-tracking, but broad triggers like 'test this', 'experiment', and 'hypothesis' create minor overlap risk with closely related marketing skills, keeping it just below a 5.

4 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
alirezarezvani/claude-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.