CtrlK
BlogDocsLog inGet started
Tessl Logo

ab-testing

When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program.

63

Quality

75%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/antigravity-ab-testing/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, actionable skill body with concrete templates, tables, named tools, and a clear experiment loop supported by checklists. The main drags are over-explanation of basic statistics Claude already knows and a somewhat long inline growth-program section that could live in a reference.

Suggestions

Trim or move explanations of basic statistics (the p-value/95%-confidence definition and the 'Peeking Problem' paragraph) since Claude already knows these; keep only the operational guidance.

Extract the large 'Growth Experimentation Program' section (experiment loop, ICE, velocity, playbook, cadence) into a references file (e.g. references/experiment-program.md) and link to it from SKILL.md to tighten the overview.

Add an explicit validate-and-retry checkpoint to the test-running workflow (e.g. 'if tracking/QA fails → fix and re-verify before launching') to lift workflow clarity toward fully explicit feedback loops.

DimensionReasoningScore

Conciseness

Mostly efficient with useful reference tables and templates, but it explains concepts Claude already knows (e.g. '95% confidence = p-value < 0.05. Means <5% chance result is random' and 'The Peeking Problem… leads to false positives'), so it is not lean enough for the 4 anchor.

3 / 5

Actionability

Provides concrete, copy-paste-ready templates (hypothesis framework, experiment playbook), specific quick-reference tables, named tools (PostHog, Optimizely, LaunchDarkly), and the ICE formula; minor gaps keep it just below fully executable at 5.

4 / 5

Workflow Clarity

The experiment loop is clearly sequenced (1–6) with pre-launch and analysis checklists plus a 'mixed signals → dig deeper' feedback path, but validation checkpoints are implicit rather than fully explicit error-recovery loops, matching the 4 anchor.

4 / 5

Progressive Disclosure

Two real one-level-deep references (sample-size-guide.md, test-templates.md) are clearly signaled, but the body is long and keeps the entire growth-experimentation-program section inline rather than splitting it out, so it does not reach the cleanly-split 5 anchor.

4 / 5

Total

15

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, trigger-forward description that clearly states both what the skill does and when to use it, with concrete actions and natural keywords. Its main weakness is limited synonym coverage and minor overlap with adjacent growth/analytics skills.

DimensionReasoningScore

Specificity

Names the experimentation domain with several concrete actions ('plan, design, or implement an A/B test', 'build a growth experimentation program'), but coverage is not comprehensive — analysis/measurement actions are absent, so it sits below the 5 anchor.

4 / 5

Completeness

Explicitly answers both 'what' (plan/design/implement A/B tests, build an experimentation program) and 'when' ('When the user wants to…') with concrete trigger phrasing, matching the top anchor.

5 / 5

Trigger Term Quality

Includes natural terms a user would say ('A/B test', 'experiment', 'growth experimentation program') but is missing common synonyms like 'split test', 'multivariate test', or 'hypothesis testing' that the body itself uses, keeping it just below comprehensive.

4 / 5

Distinctiveness Conflict Risk

Has a clear A/B-testing niche with distinct triggers, but the broad term 'experiment' and the skill's own related-skills list (cro, analytics) introduce minor overlap risk, so it does not fully reach the minimal-conflict 5 anchor.

4 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
boisenoise/skills-collections
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.