CtrlK
BlogDocsLog inGet started
Tessl Logo

experiment-design

A discipline for designing experiments (A/B tests, multivariate, holdouts) so the results actually answer the question you asked. Hypothesis writing, sample size, duration, segment analysis, running discipline, matching a result to a pre-committed decision rule, and the common failure modes that produce confidently wrong shipping decisions. Use this skill whenever the user is planning a test that has not run yet: framing a hypothesis, sizing the sample, setting duration, choosing guardrails, or deciding whether something is worth testing at all. Triggers on design an experiment, experiment plan, A/B test, split test, multivariate test, holdout, experiment hypothesis, sample size, minimum detectable effect, MDE, test duration, guardrail metric, no peeking, pre-committed decision rule, is this worth testing. Use `experimentation-analytics` instead when the test has already run and the question is how to read the result panel.

72

Quality

89%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, expert instruction-only skill: highly actionable specific guidance and excellent progressive disclosure with seven real, well-signaled reference files. The main weakness is conciseness — recurring philosophical framing and re-explanation of statistical concepts Claude already knows add tokens that could be trimmed.

Suggestions

Tighten the intro and 'Closing: when in doubt' sections: drop the rhetorical framing about the default state of experimentation and keep only the decision-relevant guidance, cutting roughly 15-20% of tokens.

Compress the re-explanations of known statistical concepts (multiple-comparisons false-positive math, peeking inflation percentages, novelty/primacy definitions) into one-line reminders that point to the decision rule rather than re-deriving them.

Promote the pre-experiment readiness and post-experiment decision steps into a single explicit numbered workflow with validation checkpoints in SKILL.md, so the lifecycle sequence is unambiguous rather than implied across sections.

DimensionReasoningScore

Conciseness

The body is dense with genuinely expert, specific guidance, but it also re-explains concepts Claude already knows (the multiple-commissions math, peeking false-positive inflation, novelty/primacy effects) and includes philosophical framing ('The default state of experimentation in most companies is sloppy', the 'Closing: when in doubt' section) that could be trimmed without losing actionability.

3 / 5

Actionability

Highly concrete and specific for an instruction-only skill: 'Pick exactly one primary metric. Pick three to five guardrails', 'Two weeks is the conventional minimum for any UI/UX experiment', 'Maximum duration. Usually four to six weeks', and the vendor question 'What is your variance estimator for ratio metrics?' — absence of code is not penalized because the guidance is directly actionable.

5 / 5

Workflow Clarity

The 12-consideration framework plus the lifecycle ordering (readiness -> hypothesis -> sizing -> duration -> running -> interpretation -> decision) and ranked inconclusive-resolution paths give a clear sequence with checkpoints (falsifiability test, pre-commitment, referenced readiness checklist), though it reads more as a playbook than a strictly validated step sequence.

4 / 5

Progressive Disclosure

Well-structured overview with one-level-deep references to seven verified bundle files (hypothesis-templates, sample-size-tables, common-failures, results-interpretation-checklist, platform-comparison, pre-experiment-readiness-checklist, post-experiment-decision-framework), each clearly signaled via markdown links and a dedicated 'Reference files' section with descriptions.

5 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is exemplary: third-person voice, comprehensive concrete capabilities, rich natural trigger terms, explicit use-when guidance, and a clear disambiguation pointer to a related skill. It fully satisfies all four dimensions.

DimensionReasoningScore

Specificity

Lists multiple concrete capabilities — 'Hypothesis writing, sample size, duration, segment analysis, running discipline, matching a result to a pre-committed decision rule' plus 'choosing guardrails, or deciding whether something is worth testing at all' — giving comprehensive coverage of the experiment-design lifecycle.

5 / 5

Completeness

Explicitly answers both what (the design discipline and its components) and when ('Use this skill whenever the user is planning a test that has not run yet: framing a hypothesis, sizing the sample, setting duration, choosing guardrails...'), with concrete trigger phrases.

5 / 5

Trigger Term Quality

Comprehensive natural-term coverage including synonyms and abbreviations: 'A/B test, split test, multivariate test, holdout, experiment hypothesis, sample size, minimum detectable effect, MDE, test duration, guardrail metric, no peeking, pre-committed decision rule, is this worth testing'.

5 / 5

Distinctiveness Conflict Risk

Clear niche (pre-test design) with an explicit boundary to a sibling skill — 'Use experimentation-analytics instead when the test has already run and the question is how to read the result panel' — minimizing wrong-skill triggering.

5 / 5

Total

20

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
rampstackco/claude-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.