CtrlK
BlogDocsLog inGet started
Tessl Logo

analyzing-experiment-session-replays

Analyze session replay patterns across experiment variants to understand user behavior differences. Use when the user wants to see how users interact with different experiment variants, identify usability issues, compare behavior patterns between control and test groups, or get qualitative insights to complement quantitative experiment results.

75

Quality

92%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

85%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, highly actionable workflow skill with concrete SQL and filter examples plus solid validation checkpoints. Its main weakness is redundancy: the example interaction and notes sections repeat details already present in the workflow steps.

Suggestions

Consolidate filter-construction guidance: the 'Key points', 'Important notes', and 'Example interaction' sections restate the $feature/<flag_key> filter and boolean-flag value handling already detailed in step 2 — keep it in one place and cross-reference.

Trim the worked 'Example interaction' block or convert it to a brief inline reference, since it largely re-walks the 5-step workflow with sample numbers rather than adding new guidance.

Move the 'Related tools' list nearer the workflow steps (or into step 1/3 where each tool is introduced) so tool selection guidance is colocated with usage.

DimensionReasoningScore

Conciseness

The body is mostly efficient and domain-specific (e.g. the $feature/<flag_key> property is genuinely non-obvious), but the 'Example interaction', 'Important notes', and 'Key points' sections restate filter-construction and verification details already covered in the workflow, so it could be tightened.

2 / 3

Actionability

It provides fully executable guidance: concrete HogQL queries, a copy-paste-ready JSON filter structure with exact field values, named tools (query-session-recordings-list, experiment-get), and a worked example — matching the 'fully executable, copy-paste ready' anchor.

3 / 3

Workflow Clarity

A clearly sequenced 5-step process with explicit validation checkpoints — verify the experiment is launched, confirm recordings exist before analyzing, and verify the flag filter actually filters (nonexistent key should return zero recordings) — plus error-handling feedback for draft state and missing recordings.

3 / 3

Progressive Disclosure

No bundle files exist; the single SKILL.md is well-organized into logical sections (When to use, Prerequisites, Workflow, Example, Important notes, Related tools) with no nested or multi-level references, so the content is appropriately self-contained and easy to navigate.

3 / 3

Total

11

/

12

Passed

Description

100%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, third-person description that names concrete capabilities, includes an explicit 'Use when' trigger clause with natural user phrasing, and occupies a clear niche unlikely to conflict with other skills. No vague fluff or over-claims.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — 'Analyze session replay patterns', 'identify usability issues', 'compare behavior patterns between control and test groups', 'get qualitative insights' — rather than vague language, matching the 'lists multiple specific concrete actions' anchor.

3 / 3

Completeness

It clearly states what the skill does in the first sentence and provides an explicit 'Use when the user wants to...' clause with concrete trigger conditions, satisfying both 'what' and 'when'.

3 / 3

Trigger Term Quality

It covers natural terms a user would say — 'session replays', 'experiment variants', 'control and test groups', and quotes a likely user phrasing ('how users interact with different experiment variants') — giving good coverage of common variations.

3 / 3

Distinctiveness Conflict Risk

The niche — analyzing session replays scoped to experiment variants — is distinct, with triggers ('experiment variants', 'control and test groups') unlikely to fire for unrelated skills.

3 / 3

Total

12

/

12

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
PostHog/posthog
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.