CtrlK
BlogDocsLog inGet started
Tessl Logo

review-experiment

Use when reviewing experiment results against a declared hypothesis, primary metric, and decision rule.

70

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Review Experiment

Purpose

Produce a read-only review preserving the declared test and limiting conclusions to supplied evidence.

Required inputs

  • Predeclared hypothesis, variants, and primary metric
  • Metric numerator, denominator, analysis unit, and attribution window
  • Sample counts, eligibility or exclusions, and assignment method per variant
  • Outcomes and observation period
  • Predeclared decision rule, sample plan, and analysis method
  • Known changes or confounders

Label absent or conflicting inputs Missing; never reconstruct them after results.

Workflow

  1. Restate the hypothesis, variants, and primary metric exactly as declared. Do not switch outcomes, segments, windows, or thresholds after inspecting results.
  2. Define the primary metric, unit, and attribution window. If missing, report arithmetic only on supplied records; do not relabel records as sent, delivered, exposed, unique prospects, or attributable replies.
  3. Report counts, exclusions, assignment, and balance per variant. Distinguish randomized, matched, and unknown assignment; never assume causal design.
  4. Reproduce supplied outcomes and show compatible arithmetic with formulas. Label it Descriptive arithmetic, not lift, significance, effectiveness, equivalence, or causation.
  5. Separate Supplied confounders from Unmeasured possibilities. Possibilities are not facts or explanations.
  6. Apply only the predeclared rule and method. Describe uncertainty and count-based sample sensitivity; do not invent formal power, a p-value, interval, threshold, winner, or significance. Equal rates do not prove equivalence.
  7. Return INCONCLUSIVE whenever required definitions, assignment evidence, a valid comparison, the predeclared decision rule, or sufficient evidence under that rule is missing. Otherwise report the rule-derived decision without causal language beyond the design.
  8. Draft the smallest next step, such as fixing measurement, predeclaring a rule, or collecting the planned sample. Take no external action.

Output format

# Experiment review draft
- Status: DRAFT ONLY — READ-ONLY REVIEW
- Decision: [rule-derived decision or INCONCLUSIVE]
## Hypothesis
- [declared hypothesis and variants]
## Primary metric
- [definition, unit, window, or Missing]
## Sample and assignment
- [counts, exclusions, assignment, balance, or Missing]
## Supplied result
- [outcomes and descriptive arithmetic with formula]
## Confounders
- Supplied: [items or None supplied]
- Unmeasured possibilities: [clearly labeled items]
## Decision rule
- [rule, application, sensitivity, and missing evidence]
## Draft recommendation
- [reviewable next step]
## Actions taken
- None

Guardrails

  • Do not invent evidence, significance, lift, causation, equivalence, a winner, delivery, or system state.
  • Reject targeting, segmentation, or decisions based on protected traits (such as race, religion, disability, sex, or age) and use of unnecessary sensitive personal data.
  • Do not browse, contact recipients, alter an experiment, or mutate any external system.

Quality check

Confirm hypothesis and metric stayed fixed; inputs, result, confounders, and rule are explicit; insufficient evidence yields INCONCLUSIVE; and no action occurred.

Example invocation

Use review-experiment on a source-cited-opener hypothesis with 10 offline records and one supplied positive outcome per variant; keep missing assignment and decision-rule evidence explicit.

Repository
llaskin/AI-SDR-Skill-Pack
Last updated
First committed

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.