CtrlK
BlogDocsLog inGet started
Tessl Logo

ab-testing

When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program.

50

Quality

55%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/ab-testing/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A comprehensive A/B testing skill with strong actionability through concrete frameworks, templates, and reference tables. The main weaknesses are verbosity (explaining concepts Claude already knows like statistical significance basics) and length that would benefit from better progressive disclosure into supporting files. The workflow is well-structured with good checklists and validation steps, though the referenced bundle files don't exist.

Suggestions

Trim explanations of concepts Claude already knows (statistical significance definition, what A/B testing is, the peeking problem explanation) — a brief reminder is sufficient

Move the Growth Experimentation Program section, sample size reference tables, and variant design guidance into the referenced supporting files to reduce SKILL.md length

Provide the referenced bundle files (references/sample-size-guide.md and references/test-templates.md) or remove the references to avoid broken links

Add a single end-to-end workflow summary at the top that sequences the full process (assess → hypothesize → design → implement → run → analyze → document) before diving into section details

DimensionReasoningScore

Conciseness

The skill contains useful reference tables and frameworks, but is notably verbose at ~300+ lines. Several sections explain concepts Claude already knows (what statistical significance means, what A/B testing is, the peeking problem). The 'Common Mistakes' section largely restates advice already given. The hypothesis framework explanation and test types table add value but could be tighter.

3 / 5

Actionability

Provides concrete frameworks (hypothesis template, ICE scoring, sample size tables, checklists, experiment playbook template) that are directly usable. However, there's no executable code — the skill is instruction-oriented, which is appropriate for the domain. Minor gaps: the checklist items are good but some guidance remains high-level (e.g., 'tracking verified' without specifying how).

4 / 5

Workflow Clarity

The experiment loop is clearly sequenced, the pre-launch checklist provides validation steps, and the analysis checklist gives a clear post-test workflow. The cadence section adds temporal structure. Minor gap: the overall flow from initial assessment through to documentation could be more explicitly sequenced as a single end-to-end workflow rather than presented as separate sections.

4 / 5

Progressive Disclosure

References two external files (references/sample-size-guide.md and references/test-templates.md) which is good structure, but no bundle files are provided so these are broken references. The skill itself is quite long and could benefit from moving the Growth Experimentation Program section, sample size tables, and variant design guidance into separate reference files. The product-marketing context file check is a nice touch.

3 / 5

Total

14

/

20

Passed

Description

47%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description functions primarily as a trigger clause without explaining what the skill actually does. While it includes decent trigger terms for A/B testing scenarios, it completely lacks specificity about the skill's capabilities — does it generate test plans, calculate sample sizes, analyze results, or write code? The absence of a 'what' component significantly undermines its usefulness for skill selection.

Suggestions

Add concrete actions describing what the skill does, e.g., 'Designs experiment hypotheses, calculates sample sizes, defines success metrics, creates implementation plans, and analyzes test results for A/B tests and growth experiments.'

Include additional trigger synonyms like 'split test', 'multivariate test', 'conversion optimization', 'experiment design', or 'statistical significance' to improve matching.

Restructure to lead with capabilities (what it does) followed by the trigger clause (when to use it), e.g., 'Designs and plans A/B tests including hypothesis formation, sample size calculation, and metric definition. Use when the user wants to plan, design, or implement an A/B test or experiment.'

DimensionReasoningScore

Specificity

Names the domain (A/B testing, experimentation) but provides no concrete actions. 'Plan, design, or implement' are generic verbs that don't describe specific capabilities like statistical analysis, sample size calculation, or variant creation.

2 / 5

Completeness

Has a 'when' clause ('When the user wants to...') but the 'what' is essentially absent — it doesn't describe what the skill actually does or what capabilities it provides. The description only specifies trigger conditions without explaining the skill's concrete outputs or actions.

2 / 5

Trigger Term Quality

Includes good natural trigger terms like 'A/B test', 'experiment', and 'growth experimentation program' that users would naturally say. Missing some common synonyms like 'split test', 'multivariate test', 'feature flag', 'conversion optimization', or 'hypothesis testing'.

4 / 5

Distinctiveness Conflict Risk

The A/B testing and experimentation domain is fairly distinct and unlikely to conflict with most other skills. Minor overlap risk with general data analysis or statistics skills, but the specific mention of 'A/B test' and 'growth experimentation program' provides reasonable differentiation.

4 / 5

Total

12

/

20

Passed

Validation

90%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation10 / 11 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

10

/

11

Passed

Repository
administrakt0r/AI-Agents-Safe-Coding-Skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.