CtrlK
BlogDocsLog inGet started
Tessl Logo

ab-test-setup

Structured guide for setting up A/B tests with mandatory gates for hypothesis, metrics, and execution readiness.

57

Quality

65%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/ab-test-setup/SKILL.md

The canonical home for this skill is ab-test-setup in sickn33/agentic-awesome-skills

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is a well-structured, gated A/B-test workflow with strong validation checkpoints and concrete checklists, scoring high on workflow clarity and actionability. Its main weakness is progressive disclosure: existing reference files are not linked from the body, leaving navigation gaps and inlined reference content.

Suggestions

Link the existing references from the body, e.g. in §7 'See references/sample-size-guide.md for quick-reference tables and the duration formula' and in Documentation 'Use references/test-templates.md for the test plan and results templates'.

Move the sample-size input definitions out of the body and point to sample-size-guide.md to reduce inlining and respect token budget.

Trim redundant reinforcement sections ('Key Principles', 'Final Reminder') that restate the earlier gates to improve conciseness.

DimensionReasoningScore

Conciseness

The body is mostly efficient with tight bullet lists and concrete gates, but sections like 'Key Principles (Non-Negotiable)' and the motivational 'Final Reminder' restate earlier content and could be trimmed.

4 / 5

Actionability

Concrete gating questions, a 5-step tracking-verification checklist with specific thresholds (±5%, within 30 seconds, 5+ events), and a results decision table give mostly executable guidance; a few sections remain conceptual checklists rather than worked examples.

4 / 5

Workflow Clarity

A clearly numbered gate sequence (1-8) with explicit hard stops, a dedicated pre-Gate-8 verification checklist, and feedback loops ('stop and resolve it') match the anchor for clear sequencing with explicit validation and error-recovery checklists.

5 / 5

Progressive Disclosure

Two bundle files exist (test-templates.md, sample-size-guide.md) that map directly onto the sample-size and documentation sections, but the body never links or signals them, and content that belongs in those references is inlined.

3 / 5

Total

16

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description clearly states what the skill does and names a specific domain with concrete gating mechanisms, but it omits any explicit "when to use" trigger guidance and lacks keyword synonyms. Adding a trigger clause with natural phrases would lift completeness and trigger_term_quality.

Suggestions

Append a 'Use when...' clause with natural trigger phrases, e.g. 'Use when setting up A/B tests, split tests, or experiments, or when the user asks about test hypothesis, sample size, or execution readiness.'

Add synonym keywords (split test, experiment, conversion-rate optimization) to broaden trigger coverage.

Enumerate 2-3 more concrete actions (e.g. 'lock hypothesis, calculate sample size, verify tracking') to raise specificity.

DimensionReasoningScore

Specificity

Names the A/B-test domain and a few concrete mechanisms ("mandatory gates for hypothesis, metrics, and execution readiness") but does not enumerate several specific concrete actions, so it is not comprehensive enough for a 4.

3 / 5

Completeness

It gives a clear "what" (a structured gated guide for A/B test setup) but provides no "Use when..." clause or equivalent trigger guidance, which caps completeness at 3 per the rubric guidelines.

3 / 5

Trigger Term Quality

"A/B tests" is a natural, relevant keyword, but synonyms and variations users say (split test, experiment, conversion optimization) are missing and there is no trigger clause.

3 / 5

Distinctiveness Conflict Risk

"A/B test setup" is a mostly-distinct niche unlikely to trigger for unrelated skills, though it could overlap slightly with general experimentation/analysis skills, keeping it below 5.

4 / 5

Total

13

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
administrakt0r/AI-Agents-Safe-Coding-Skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.