CtrlK
BlogDocsLog inGet started
Tessl Logo

testland/ab-test-validity-checklist

Workflow skill that builds an A/B-test validity checklist from an experiment proposal, walking the canonical design-correctness gates - pre-registered OEC/power/guardrails, randomization unit + SRM check, assignment integrity, telemetry, peeking discipline, novelty/primacy, post-experiment SRM re-check - into a per-experiment checklist + sign-off form. Use when launching, auditing, or governing an experiment. For pitfall mechanics use guardrail-metrics-reference or peeking-problem-reference; to read an already-valid result use experiment-results-interpreter; for per-SDK harness tests use optimizely-test or statsig-test - this gates DESIGN, not SDK code.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Overview
Quality
Evals
Security
Files

Quality

Content

85%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a highly actionable, well-sequenced gated checklist with executable thresholds and a ready-to-emit template. It is slightly token-heavy in restating experiment-statistics domain knowledge and inlines reference-style material that would benefit from separate companion files.

Suggestions

Move the guardrail-class threshold catalog (Step 1) into a referenced companion file so the body signals it rather than inlining the full table, improving progressive disclosure and token efficiency.

Trim the worked SRM chi-square numeric example to a one-line threshold statement, trusting Claude's existing statistics knowledge.

Create the referenced bundle files (guardrail-metrics-reference, peeking-problem-reference) or remove the inline cross-references so signaled paths resolve to real files.

DimensionReasoningScore

Conciseness

Mostly efficient and tightly table-driven, but the SRM chi-square worked example (expected/observed/χ²=36/p-value) and the guardrail threshold catalog restate domain knowledge and consume tokens that don't all earn their place for an agent that already knows experiment statistics.

2 / 3

Actionability

Provides concrete, copy-paste-ready outputs: an executable chi-square formula, specific block thresholds (e.g. '> 10% or > 50ms', 'p < 0.0001'), and a full markdown checklist + sign-off template ready to emit.

3 / 3

Workflow Clarity

Eight numbered steps form a clear pre-flight→post-flight sequence with explicit validation gates (SRM 'invalid until root cause found', Step 7 pass criteria table, 'stop ship discussion' checkpoint) and a feedback loop, plus an Anti-patterns table covering failure recovery.

3 / 3

Progressive Disclosure

Well-organized into clear sections and signals companion catalogs inline ('per guardrail-metrics-reference', 'per peeking-problem-reference'), but no bundle files exist and the body inlines a sizable guardrail threshold table and worked SRM example that read as reference material better split into separate files.

2 / 3

Total

10

/

12

Passed

Description

100%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific, trigger-rich, and explicitly scopes when to use the skill versus when to defer to companion skills. It cleanly answers both what and when in third person with no fluff.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'builds an A/B-test validity checklist', 'walking the canonical design-correctness gates', 'pre-registered OEC/power/guardrails', 'SRM check', 'assignment integrity', 'telemetry', 'peeking discipline', 'post-experiment SRM re-check' — naming specific gates rather than abstract language.

3 / 3

Completeness

Explicitly answers both 'what' (builds a per-experiment checklist + sign-off form from the design gates) and 'when' with a clear 'Use when launching, auditing, or governing an experiment' trigger clause.

3 / 3

Trigger Term Quality

Covers natural user-facing terms ('launching, auditing, or governing an experiment', 'experiment proposal', 'A/B-test', 'validity checklist') plus concrete skill-name triggers, giving good coverage of phrasings a user would actually say.

3 / 3

Distinctiveness Conflict Risk

Carves a clear niche — 'this gates DESIGN, not SDK code' — and explicitly routes adjacent intents to other skills (guardrail-metrics-reference, peeking-problem-reference, experiment-results-interpreter, optimizely-test, statsig-test), making wrong-skill conflicts unlikely.

3 / 3

Total

12

/

12

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Reviewed

Table of Contents