CtrlK
BlogDocsLog inGet started
Tessl Logo

ab-testing

When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "should I test this," "which version is better," "test two versions," "statistical significance," "how long should I run this test," "growth experiments," "experiment velocity," "experiment backlog," "ICE score," "experimentation program," or "experiment playbook." Use this whenever someone is comparing two approaches and wants to measure which performs better, or when they want to build a systematic experimentation practice. For tracking implementation, see analytics. For page-level conversion optimization, see cro.

69

Quality

84%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

73%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, actionable methodology skill with clear workflows, checklists, and properly signaled one-level references. The main drag is conciseness: several sections explain statistics concepts Claude already knows, and the body could offload more material to reference files.

Suggestions

Trim or remove explanations of well-known concepts (statistical significance definition, the peeking problem, what guardrail/secondary metrics are, ICE) since Claude already knows them; keep only the skill-specific framing.

Consider moving the large Growth Experimentation Program section (experiment loop, ICE, velocity, playbook, cadence) into a dedicated reference file, leaving a short overview and a clearly signaled link in SKILL.md.

Add a brief inline note on how to actually compute significance/confidence intervals (or a one-line pointer to the specific section of the sample-size reference) so the analysis step is fully executable without relying on external calculators.

DimensionReasoningScore

Conciseness

Largely efficient through tables, checklists, and templates, but several sections restate concepts Claude already knows (the definition of statistical significance, the peeking problem, what guardrail metrics are, ICE scoring) that could be tightened or trimmed.

3 / 5

Actionability

Provides concrete, copy-paste-ready templates (hypothesis structure, experiment playbook), quick-reference tables (sample size, test types, traffic allocation), and checklists; the actual statistical computation steps are delegated to calculators and a reference file, leaving minor gaps.

4 / 5

Workflow Clarity

Clear end-to-end sequence (assess → hypothesis → design → implement → run → analyze → document, plus the experiment loop) with explicit validation checkpoints (pre-launch checklist with tracking/QA verification) and feedback loops (guardrail stop-rules, monthly re-ICE).

5 / 5

Progressive Disclosure

Two real, well-signaled one-level-deep references (sample-size-guide.md, test-templates.md) with the body acting as overview; however the body is long and the sizable Growth Experimentation Program section could arguably be split into its own reference.

4 / 5

Total

16

/

20

Passed

Description

95%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that clearly states capabilities, provides extensive natural trigger terms, and explicitly distinguishes itself from adjacent skills. The only minor weakness is that the named actions are relatively high-level rather than granular.

DimensionReasoningScore

Specificity

Lists several concrete actions ("plan, design, or implement an A/B test or experiment, or build a growth experimentation program") but the verbs are somewhat high-level compared to granular task-level actions, leaving minor coverage gaps.

4 / 5

Completeness

Explicitly answers both what (plan/design/implement A/B tests, build an experimentation program) and when (multiple "Use when..." / "Also use when the user mentions..." / "Use this whenever..." clauses) with concrete trigger phrases, in third person.

5 / 5

Trigger Term Quality

Comprehensive coverage of natural trigger terms including synonyms and phrasings users actually say ("A/B test," "split test," "multivariate test," "statistical significance," "ICE score," "how long should I run this test").

5 / 5

Distinctiveness Conflict Risk

Occupies a clear niche with distinct triggers and explicitly disambiguates sibling skills ("For tracking implementation, see analytics. For page-level conversion optimization, see cro."), minimizing conflict risk.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
joshmanders/dotfiles
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.