CtrlK
BlogDocsLog inGet started
Tessl Logo

ab-testing

When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "should I test this," "which version is better," "test two versions," "statistical significance," "how long should I run this test," "growth experiments," "experiment velocity," "experiment backlog," "ICE score," "experimentation program," or "experiment playbook." Use this whenever someone is comparing two approaches and wants to measure which performs better, or when they want to build a systematic experimentation practice. For tracking implementation, see analytics. For page-level conversion optimization, see cro.

68

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, actionable skill body with strong templates, tables, and checklists, and correctly offloaded detailed material into two real reference files. Its main weakness is conciseness: it re-explains basic statistical concepts Claude already knows.

Suggestions

Trim or remove the 'The Peeking Problem' paragraph and the statistical-significance basics ('95% confidence = p-value < 0.05 ... not a guarantee'); Claude already knows these — keep only the operational rule (pre-commit to sample size, don't stop early).

Tighten the 'Common Mistakes' section into a compact checklist or fold it into the existing pre-launch and analysis checklists to avoid restating obvious pitfalls.

Consider moving the full experiment-playbook template block into references/test-templates.md and linking to it, leaving only a short inline example in SKILL.md to improve progressive disclosure.

DimensionReasoningScore

Conciseness

Mostly efficient tables and checklists, but several sections re-explain concepts Claude already knows (e.g., 'The Peeking Problem ... leads to false positives', '95% confidence = p-value < 0.05 ... Not a guarantee—just a threshold'), fitting the anchor-3 'mostly efficient but includes some unnecessary explanation'.

3 / 5

Actionability

Provides concrete, copy-paste-ready templates (hypothesis framework, experiment playbook) and specific quick-reference tables (sample size, ICE, traffic allocation); the minor gap is reliance on external calculators rather than an executable command for sample sizing.

4 / 5

Workflow Clarity

A clear numbered Experiment Loop plus pre-launch and analysis checklists give a well-sequenced workflow with most checkpoints present; the 'During the Test' monitoring guidance is less concrete about recovery actions, a minor validation gap.

4 / 5

Progressive Disclosure

SKILL.md acts as an overview with well-signaled, one-level-deep links to the real bundle files (references/sample-size-guide.md, references/test-templates.md) for detailed tables and templates; structure is good, though a fair amount of inlined material (full playbook template, cadence) is borderline bulk that could live in references.

4 / 5

Total

15

/

20

Passed

Description

95%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, comprehensive description that clearly states capabilities, enumerates extensive natural trigger phrases, and disambiguates from sibling skills. The only minor weakness is that the named actions are relatively high-level verbs rather than highly granular tasks.

DimensionReasoningScore

Specificity

Lists several concrete actions ('plan, design, or implement an A/B test or experiment, or build a growth experimentation program') but the verbs are somewhat high-level rather than granular like the anchor-5 PDF example, leaving minor coverage gaps.

4 / 5

Completeness

Explicitly answers both what ('plan, design, or implement an A/B test ... or build a growth experimentation program') and when ('When the user wants to ...', 'Also use when the user mentions ...', 'Use this whenever someone is comparing two approaches') with concrete trigger phrases.

5 / 5

Trigger Term Quality

Comprehensive coverage of natural user phrases including synonyms and question forms ('A/B test,' 'split test,' 'multivariate test,' 'should I test this,' 'which version is better,' 'how long should I run this test,' 'ICE score'), matching the anchor-5 expectation.

5 / 5

Distinctiveness Conflict Risk

Clear experimentation niche with explicit boundary guidance ('For tracking implementation, see analytics. For page-level conversion optimization, see cro.'), minimizing overlap with related skills.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
coreyhaines31/marketingskills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.