CtrlK
BlogDocsLog inGet started
Tessl Logo

configuring-experiment-analytics

Configures the analytics side of a PostHog experiment — exposure criteria (server-resolved default exposure event vs custom exposure events), primary and secondary metrics, the supported metric types (count, sum, ratio with `math` and `math_property`, retention with `retention_window_start` and `start_handling`), multivariate user handling ("Exclude" vs "First seen variant"), and how to read results once the experiment is live. Use when the user adds or edits a primary or secondary metric (e.g. "add a secondary metric tracking 'downloaded_file' per user"), sets up a ratio metric (e.g. "revenue from purchase_completed / pageviews"), sets up a retention metric (e.g. "$pageview → uploaded_file, 7-day window"), configures custom exposure (e.g. "only count users who hit /checkout"), changes multivariate handling, or asks "who is in the analysis?", "how do I measure impact?", "is this winning?", "what's the confidence level?", or "should I ship?".

69

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

—

The risk profile of this skill

SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, actionable body with concrete API calls, field shapes, and a clearly sequenced metric-building workflow guarded by pre-flight validation. Its main defect is that the three `references/*.md` files it points to for full schemas, templates, and result interpretation are absent, breaking the progressive-disclosure path.

Suggestions

Add the missing `references/` bundle files (`metric-templates.md`, `metric-configuration.md`, `interpreting-results.md`) the body already cites, or inline the critical bits and remove the dangling references.

Include at least one complete copy-paste-ready inline `ExperimentMetric` JSON payload (e.g. a ratio metric) so the skill is actionable even before the reference file is consulted.

Add an explicit post-update verification step (e.g. re-call `experiment-get` to confirm the metric attached with the intended type) to close the destructive replace-list workflow with a feedback loop.

DimensionReasoningScore

Conciseness

Mostly efficient and product-specific (assumes Claude knows what APIs/flags are), but the WRONG/RIGHT example block and the extended bias-risk prose could be trimmed without losing the operational point.

4 / 5

Actionability

Highly concrete — names exact endpoints (`experiment-saved-metrics-list?event=`, `experiment-update`), field shapes, and parameter names (`saved_metrics_ids`, `math: "sum"`, `allow_unknown_events`), but defers full inline JSON payloads to a reference file rather than giving copy-paste-ready payloads in the body.

4 / 5

Workflow Clarity

Clear Step 1→4 sequence with pre-flight validation checkpoints ('always get the current experiment first via experiment-get', 'MUST call read-data-schema', 'confirm the match with the user') for the destructive list-replacing operations; minor gap is the absence of an explicit post-update verification step.

4 / 5

Progressive Disclosure

Structure and signaling are good — clear sections and one-level-deep references to `references/metric-templates.md`, `references/metric-configuration.md`, and `references/interpreting-results.md` — but the `references/` directory does not exist, so those signaled references dangle and the disclosure is non-functional.

3 / 5

Total

15

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: concrete, comprehensive, third-person, with an explicit 'Use when...' clause packed with natural trigger phrases and worked examples. It cleanly answers what the skill does and when to invoke it with minimal conflict risk.

DimensionReasoningScore

Specificity

Lists multiple concrete capabilities — exposure criteria, primary/secondary metrics, the four metric types with their specific fields (`math`, `math_property`, `retention_window_start`, `start_handling`), multivariate handling, and result interpretation — giving comprehensive coverage of the analytics side of an experiment.

5 / 5

Completeness

Explicitly answers both 'what' (configures exposure, metrics, metric types, multivariate handling, results) and 'when' via a detailed 'Use when...' clause with concrete trigger phrases and worked examples.

5 / 5

Trigger Term Quality

The 'Use when...' clause surfaces natural user phrasings a person would actually say — 'add a secondary metric tracking downloaded_file per user', 'revenue from purchase_completed / pageviews', 'is this winning?', 'should I ship?', 'what's the confidence level?' — with strong synonym/variation coverage.

5 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (PostHog experiment analytics) with distinct, domain-specific triggers that are unlikely to fire for the sibling rollout/diagnosing skills it itself names as related, minimizing conflict risk.

5 / 5

Total

20

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 3 missing

Warning

Total

15

/

16

Passed

Repository
PostHog/posthog
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.