CtrlK
BlogDocsLog inGet started
Tessl Logo

ab-test-setup

Design, analyze, and document A/B tests for conversion, onboarding, pricing, lifecycle, and product experiments. Use when the user asks for `/ab-test-setup`, experiment design, sample size, statistical significance, A/B test analysis, ICE-scored test backlogs, or avoiding common testing mistakes.

73

Quality

90%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

80%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-organized, concise instruction skill with concrete templates and examples. Its main weakness is workflow clarity: the experiment lifecycle and ship/revert decisions lack explicit validation checkpoints and feedback loops.

Suggestions

Add an explicit validation checkpoint before analysis (e.g., 'Confirm the test has reached the agreed sample size AND no guardrail metric was breached before interpreting results').

Provide the actual sample-size formula or a named method (e.g., two-proportion z-test power calculation) rather than only stating 'Estimate sample size'.

Include an error-recovery feedback loop for the analyze/decide step (e.g., if guardrails are violated → revert; if underpowered → continue testing rather than ship).

DimensionReasoningScore

Conciseness

The body is lean — a tight Overview, an 8-step workflow, the ICE pattern, and concrete example prompts — assuming Claude's competence without explaining statistics concepts it already knows.

5 / 5

Actionability

Provides a concrete hypothesis template, ICE definitions, and implementation checklist items, but 'Estimate sample size' omits the actual formula/method, leaving a minor gap in executable detail.

4 / 5

Workflow Clarity

An 8-step sequence is present with pre-launch decision rules, but validation checkpoints for consequential ship/revert decisions are only implicit (e.g., 'after the test reaches the agreed sample size') with no error-recovery feedback loop.

3 / 5

Progressive Disclosure

Under 50 lines with no need for external references, organized into clear sections (Overview/Workflow/Backlog/Examples), satisfying the simple-skill exception for a top score.

5 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is exemplary: it states concrete capabilities, provides comprehensive natural trigger terms, answers both what and when explicitly, and occupies a distinct niche. Third-person voice is used correctly with no over-claims.

DimensionReasoningScore

Specificity

Lists multiple concrete actions ('Design, analyze, and document A/B tests') across named experiment domains (conversion, onboarding, pricing, lifecycle) with comprehensive coverage, matching the score-5 anchor.

5 / 5

Completeness

Explicitly answers both 'what' (design/analyze/document A/B tests) and 'when' via a 'Use when...' clause with concrete trigger phrases, matching the score-5 anchor.

5 / 5

Trigger Term Quality

Comprehensive natural trigger phrases ('experiment design, sample size, statistical significance, A/B test analysis, ICE-scored test backlogs') plus the slash command — exactly what users say, with synonyms included.

5 / 5

Distinctiveness Conflict Risk

Clear A/B-testing niche with distinct triggers (/ab-test-setup, statistical significance, ICE backlogs) and minimal overlap risk with other skills.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
eigent-ai/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.