CtrlK
BlogDocsLog inGet started
Tessl Logo

experimentation-analytics

How to read experiment results without fooling yourself. Confidence intervals, p-values, multiple testing, sequential testing, CUPED, heterogeneous treatment effects, ratio metrics, network effects, dashboard reconciliation, and the interpretation failures that produce confidently wrong shipping decisions. Use this skill whenever the user is reading a finished experiment result panel and about to make a ship, kill, or iterate decision, or when an experiment number does not match the dashboard number. Triggers on read experiment results, result panel, ship or kill decision, p-value, confidence interval, statistical significance, multiple testing, peeking, sequential testing, CUPED, variance reduction, heterogeneous treatment effects, ratio metric, network effects, inconclusive test, experiment versus dashboard mismatch. Use `experiment-design` instead when the test has not run yet and the question is hypothesis, sample size, duration, or what to test.

66

Quality

80%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/experimentation-analytics/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

65%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-structured interpretation playbook with concrete decision rules and excellent progressive disclosure to seven real reference files. It loses points for re-teaching statistical basics Claude already knows and for being a reference catalog rather than a checkpointed workflow.

Suggestions

Trim or relocate definitions of concepts Claude already knows (p-value, CI, SUTVA, CUPED acronym, Bonferroni/BH) into the reference files, keeping only the interpretation discipline inline.

Add an explicit short decision workflow at the top (e.g. check panel completeness -> read CI -> check guardrails -> decide ship/kill/iterate) with validation checkpoints, rather than only a 14-item reference framework.

Move the per-platform support callouts (Statsig, Eppo, PostHog, etc.) into analytics-platform-comparison.md to reduce repetition across the CUPED, sequential-testing, and ratio-metric sections.

DimensionReasoningScore

Conciseness

The body is mostly efficient interpretation guidance but re-explains statistical concepts Claude already knows (the definition of a p-value, what a 95% CI means, the SUTVA acronym, CUPED's expansion, Bonferroni/BH), and at ~325 lines could be tightened, matching the "mostly efficient but includes some unnecessary explanation" anchor.

3 / 5

Actionability

Though code-free, it gives concrete executable guidance: five numbered CI decision rules with explicit interval examples (e.g. [-1%, +5%]), a platform support checklist ("ask 'what variance estimator do you use for ratio metrics?'"), and a worked ratio-metric example, with only minor gaps.

4 / 5

Workflow Clarity

A 14-consideration framework and "Read the relevant section before making the decision, not after" provide a checklist sequence, but the body is a reference playbook rather than a sequenced procedure, and validation checkpoints are implicit rather than explicit.

3 / 5

Progressive Disclosure

The body keeps overview-level content inline and offloads depth to seven real, one-level-deep reference files (confirmed present in ./references/), each clearly signaled with inline contextual links and summarized in a dedicated "Reference files" section, matching the clear-overview anchor.

5 / 5

Total

15

/

20

Passed

Description

95%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is explicit, comprehensive, and well-bounded: it states what the skill covers, when to use it with concrete trigger phrases, and when to prefer a sibling skill. Its only weakness is phrasing capabilities as a topic catalog rather than action verbs, which costs it the top specificity anchor.

DimensionReasoningScore

Specificity

Names a comprehensive, concrete domain catalog ("Confidence intervals, p-values, multiple testing, sequential testing, CUPED, heterogeneous treatment effects, ratio metrics, network effects, dashboard reconciliation") plus the outcome "confidently wrong shipping decisions", but lists topics/capabilities rather than verb-driven actions, keeping it just below the action-verb anchor 5.

4 / 5

Completeness

Clearly answers both what ("How to read experiment results without fooling yourself" plus the topic list) and when ("Use this skill whenever the user is reading a finished experiment result panel and about to make a ship, kill, or iterate decision"), with concrete trigger phrases and an explicit redirect boundary.

5 / 5

Trigger Term Quality

An explicit "Triggers on" list supplies comprehensive natural phrases users would say ("ship or kill decision", "p-value", "statistical significance", "peeking", "inconclusive test", "experiment versus dashboard mismatch") including synonyms, matching the comprehensive-coverage anchor.

5 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (post-experiment result interpretation) with distinct triggers and an explicit "Use `experiment-design` instead when the test has not run yet" boundary, minimizing conflict risk.

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
rampstackco/claude-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.