CtrlK
BlogDocsLog inGet started
Tessl Logo

ab-test-setup

When the user wants to plan, design, or implement an A/B test or experiment. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," or "hypothesis." For tracking implementation, see analytics-tracking.

74

Quality

92%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

85%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, actionable experimentation guide that uses checklists, tables, and templates to drive concrete work and correctly offloads detail to two real reference files. Its main weakness is mild verbosity from restating statistical concepts Claude already knows.

Suggestions

Trim sections that re-teach known concepts (e.g. the definition of statistical significance and "The Peeking Problem") to one-line reminders, keeping only what changes Claude's behavior.

Consider moving the sample-size quick-reference table into references/sample-size-guide.md since a detailed guide already exists there, reducing inline duplication.

Tighten Core Principles truisms ("Otherwise you don't know what worked", "Not just 'let's see what happens'") into terse directives that assume Claude's competence.

DimensionReasoningScore

Conciseness

Mostly efficient tabular reference material, but it spends tokens restating concepts Claude already knows (e.g. "95% confidence = p-value < 0.05", "The Peeking Problem", and truisms like "Otherwise you don't know what worked"), so it could be tightened.

2 / 3

Actionability

As an instruction-only skill it provides concrete, copy-paste-ready assets: a fill-in hypothesis template with a worked example, a sample-size quick-reference table, named tools, and pre-launch/analysis checklists.

3 / 3

Workflow Clarity

The design→run→analyze flow is clearly sequenced, with an explicit pre-launch checklist (including "Tracking verified", "QA completed") and a six-step analysis checklist serving as validation checkpoints.

3 / 3

Progressive Disclosure

SKILL.md acts as an overview and pushes detail to two real, one-level-deep references — references/sample-size-guide.md and references/test-templates.md — each clearly signaled with context and a markdown link.

3 / 3

Total

11

/

12

Passed

Description

100%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A well-crafted description that uses third-person voice, states concrete capabilities, and provides an explicit trigger clause with natural terms. It is concise yet complete and clearly distinguishable from related skills.

DimensionReasoningScore

Specificity

Lists three concrete actions — "plan, design, or implement an A/B test or experiment" — matching the multiple-specific-actions anchor; it is not merely a domain label.

3 / 3

Completeness

Answers both what ("plan, design, or implement an A/B test or experiment") and when with an explicit "Use when the user mentions..." trigger clause, satisfying the full what-and-when anchor.

3 / 3

Trigger Term Quality

Strong natural-language coverage: "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," and "hypothesis" are all phrasings a user would actually say.

3 / 3

Distinctiveness Conflict Risk

Occupies a clear experimentation niche with distinctive triggers and even disambiguates from a sibling skill via "For tracking implementation, see analytics-tracking," reducing conflict risk.

3 / 3

Total

12

/

12

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
TheCraigHewitt/seomachine
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.