CtrlK
BlogDocsLog inGet started
Tessl Logo

diagnosing-experiment-results

Diagnoses bias, anomalies, and strange results on a PostHog experiment. Covers 0-exposure experiments, sample ratio mismatch, identity fragmentation, multi-variant exposure, uneven-split exclusion bias, significance traps (peeking, A/A, Bayesian vs Frequentist), PostHog-vs-SQL discrepancies, surprises after mid-run edits, and qualitative follow-up via a variant-split survey. TRIGGER when: user asks 'is my experiment biased?' or 'why 0 exposures?', references the bias banner, says a variant looks strange / wrong / off, sees significance flipping or A/A significance, finds PostHog numbers disagreeing with their SQL, reports surprises after mid-run edits, or wants qualitative feedback or a survey for an experiment. DO NOT TRIGGER when: creating an experiment (use creating-experiments), only configuring rollout (use configuring-experiment-rollout) or metrics (use configuring-experiment-analytics), or only asking lifecycle questions (use managing-experiment-lifecycle).

68

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

—

The risk profile of this skill

SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured routing skill with a concrete dispatch table, specific tool/field guidance, and clear sequenced steps with verification gates. Its main defect is that the body depends on seven reference files that are absent from the bundle, breaking the progressive-disclosure path and leaving the actual diagnostic procedures unreachable.

Suggestions

Ship the references/ directory: the 7 referenced files (diagnostic-snapshot.md, bias-and-skew.md, empty-experiment.md, interpretation.md, numbers-vs-sql.md, mid-run-changes.md, qualitative-feedback.md) are required for the skill to function — without them every diagnostic group routes to a missing file.

Tighten Step 3's 'don't collapse mechanisms' paragraph and Step 4's reversal-offer paragraph; the core guidance (surface co-occurring mechanisms independently; don't preemptively offer reversal) can be stated in roughly half the tokens.

Add a brief inline fallback for the most common diagnostics (e.g. the 0-exposure / SRM quick checks) so the skill remains useful even before a reference file is opened, reducing single-point dependence on the missing bundle files.

DimensionReasoningScore

Conciseness

Mostly efficient: a concrete dispatch table and a specific experiment-get field list that assume Claude's competence, but the Step 3 'don't collapse mechanisms' rationale and the Step 4 reversal-offer paragraph are lengthier than needed and could be trimmed — between anchors 3 and 5, leaning up.

4 / 5

Actionability

Gives concrete executable guidance for the routing layer ('Call experiment-get and pull these fields', a complete symptom→group table, explicit state-scoping rules), but the actual diagnostic procedures are delegated to reference files — mostly executable with the gap that deep steps live elsewhere.

4 / 5

Workflow Clarity

A clearly numbered sequence (Step 1 → 1.5 → 2 → 3 → 4) with verification gates ('verify before asking', 'only list mechanisms with a path to verification', state-calibration), but the deeper validation/feedback loops reside in the unshipped reference files, leaving minor gaps in the body itself.

4 / 5

Progressive Disclosure

The design is excellent — each diagnostic group A–F summarizes then signals a one-level-deep '→ See references/X.md' — but the bundle ships no references/ directory, so the 7 referenced detail files (diagnostic-snapshot, bias-and-skew, empty-experiment, interpretation, numbers-vs-sql, mid-run-changes, qualitative-feedback) do not exist; the overview points to materials that aren't shipped, so disclosure doesn't actually resolve.

3 / 5

Total

15

/

20

Passed

Description

95%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, third-person description that explicitly states what it does, gives concrete natural-language trigger phrases, and proactively routes away from sibling skills. The only minor weakness is limited action-verb variety relative to its rich topic enumeration.

DimensionReasoningScore

Specificity

Names the domain ('Diagnoses bias, anomalies, and strange results on a PostHog experiment') and enumerates concrete coverage (0-exposure, SRM, identity fragmentation, multi-variant exposure, uneven-split exclusion bias, significance traps, PostHog-vs-SQL, mid-run edits, qualitative survey), but the distinct action vocabulary is narrow ('Diagnoses', 'Covers', 'qualitative follow-up via survey') even though topic coverage is comprehensive — between anchors 4 and 5.

4 / 5

Completeness

Explicitly answers both what ('Diagnoses bias, anomalies... Covers [list]') and when ('TRIGGER when:' plus 'DO NOT TRIGGER when:') with concrete trigger phrases, satisfying the anchor for clearly and explicitly answering both.

5 / 5

Trigger Term Quality

The 'TRIGGER when' clause lists natural user phrasings ('is my experiment biased?', 'why 0 exposures?', 'variant looks strange / wrong / off', 'significance flipping', 'PostHog numbers disagreeing with their SQL') including synonyms — comprehensive coverage of terms a user would actually say.

5 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (PostHog experiment-result diagnosis) and actively disambiguates from sibling skills in 'DO NOT TRIGGER when' (creating-experiments, configuring-experiment-rollout/analytics, managing-experiment-lifecycle), giving minimal conflict risk.

5 / 5

Total

19

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 7 missing

Warning

referenced_paths_exist

Referenced path issues: 16 missing

Warning

Total

14

/

16

Passed

Repository
PostHog/posthog
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.