CtrlK
BlogDocsLog inGet started
Tessl Logo

ab-test-setup

Structured guide for setting up A/B tests with mandatory gates for hypothesis, metrics, and execution readiness.

48

Quality

51%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

—

The risk profile of this skill

Fix and improve this skill with Tessl

tessl review fix ./plugins/antigravity-awesome-skills/skills/ab-test-setup/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

62%Scale 1-3

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

This is a well-structured procedural skill with strong workflow clarity, featuring explicit hard gates and refusal conditions that make the process safe and bounded. Its main weaknesses are moderate verbosity (motivational content, concept explanations Claude doesn't need) and a lack of concrete examples like sample hypotheses, calculation commands, or filled-out templates. The monolithic structure is acceptable but could benefit from splitting detailed sections into referenced files.

Suggestions

Add a concrete example of a well-formed hypothesis and a poorly-formed one to make the Hypothesis Quality Checklist more actionable.

Include a specific sample size calculation formula or reference to a tool/command (e.g., a Python snippet using statsmodels) rather than just listing the required inputs.

Remove the 'Final Reminder' motivational section and trim the 'Purpose & Scope' section — Claude doesn't need encouragement to follow instructions.

Consider extracting the 'Analyzing Results' and 'Documentation & Learning' sections into separate referenced files to improve progressive disclosure.

DimensionReasoningScore

Conciseness

The skill is reasonably structured but includes some unnecessary padding — the 'Final Reminder' motivational section, the 'Purpose & Scope' preamble, and the 'When to Use' / 'Limitations' boilerplate add little value. Claude already understands A/B testing concepts like peeking and statistical power; the skill should focus on the procedural gates rather than explaining why they matter.

2 / 3

Actionability

The skill provides clear checklists and gate conditions, which are actionable for a process-oriented skill. However, it lacks concrete examples — e.g., a sample hypothesis statement, a sample size calculation formula or tool command, or a filled-out test record template. The guidance is specific enough to follow but not copy-paste ready.

2 / 3

Workflow Clarity

The multi-step process is clearly sequenced with numbered phases, two explicit hard gates (Hypothesis Lock and Execution Readiness Gate), and clear refusal conditions. The workflow includes validation checkpoints and feedback loops (e.g., 'If assumptions are weak → warn and recommend delaying'; 'If any item is missing, stop and resolve it').

3 / 3

Progressive Disclosure

The content is a monolithic single file with no references to supporting documents. At ~150 lines covering hypothesis design, metrics, sample sizing, execution, analysis, and documentation, some sections (e.g., detailed analysis discipline, documentation templates) could be split into separate referenced files. However, the internal section structure with clear headers is reasonable.

2 / 3

Total

9

/

12

Passed

Description

40%Scale 1-3

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description identifies a clear niche (A/B test setup with structured gates) which makes it distinctive, but it lacks explicit trigger guidance ('Use when...') and misses common user-facing synonyms. The actions described are limited to 'setting up' without enumerating specific capabilities.

Suggestions

Add an explicit 'Use when...' clause, e.g., 'Use when the user wants to plan, design, or set up an A/B test, split test, or experiment.'

Include common trigger term variations such as 'split test', 'experiment', 'variant testing', 'experimentation framework'.

List more specific concrete actions, e.g., 'Guides users through defining hypotheses, selecting success metrics, calculating sample sizes, and validating execution readiness before launching A/B tests.'

DimensionReasoningScore

Specificity

Names the domain (A/B tests) and mentions some specific elements (hypothesis, metrics, execution readiness), but doesn't list concrete actions beyond 'setting up'. It describes structure rather than specific capabilities like 'define hypotheses, configure metrics, validate sample sizes'.

2 / 3

Completeness

Describes what it does (structured guide for A/B test setup with gates) but has no explicit 'Use when...' clause or equivalent trigger guidance. Per the rubric, a missing 'Use when...' clause caps completeness at 2, and the 'what' is also only moderately clear, placing this at 1.

1 / 3

Trigger Term Quality

Includes 'A/B tests' which is a natural keyword, plus 'hypothesis' and 'metrics' which are relevant. However, it misses common variations like 'split test', 'experiment', 'variant testing', 'conversion testing', or 'experimentation'.

2 / 3

Distinctiveness Conflict Risk

The combination of 'A/B tests' with 'mandatory gates for hypothesis, metrics, and execution readiness' creates a clear, specific niche that is unlikely to conflict with other skills. This is a well-defined domain.

3 / 3

Total

8

/

12

Passed

Validation

90%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 10 / 11 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

10

/

11

Passed

Repository
popey/claude-code-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.