CtrlK
BlogDocsLog inGet started
Tessl Logo

analyzing-experiment-session-replays

Analyze session replay patterns across experiment variants to understand user behavior differences. Use when the user wants to see how users interact with different experiment variants, identify usability issues, compare behavior patterns between control and test groups, or get qualitative insights to complement quantitative experiment results. Also covers pairing the observed behavior with a linked survey when the user wants qualitative feedback beyond what recordings show.

65

Quality

80%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./products/experiments/skills/analyzing-experiment-session-replays/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-structured, actionable multi-step workflow with concrete queries and a worked example. Its main weakness is redundancy in the closing 'Important notes' section, which repeats earlier guidance rather than adding new value.

Suggestions

Trim the 'Important notes' section to only net-new caveats, removing the re-statements of the $feature/<flag_key> scoping and launched/start_date checks already covered in steps 1–2.

Add an explicit validation checkpoint after retrieving variants (e.g., 'If no variants returned, stop and tell the user the flag has no multivariate config') to strengthen the workflow feedback loop.

Either bring the qualitative-feedback reference into this skill's own references/ bundle or note explicitly that it is intentionally shared with diagnosing-experiment-results, so the cross-skill dependency is unambiguous.

DimensionReasoningScore

Conciseness

The body is mostly efficient with concrete queries and filters earning their place, but the 'Important notes' section restates points already made in steps 1–2 (e.g., scoping via $feature/<flag_key>, launched/start_date checks), adding redundancy that could be trimmed.

3 / 5

Actionability

Provides concrete executable HogQL queries with real table/column names, a complete filter JSON structure, named tools, and a worked example with real values; minor gaps remain in the form of placeholders like <experiment_id> and <feature_flag_key>.

4 / 5

Workflow Clarity

A clear numbered 1–6 sequence with prerequisites and a verification hint ('verify it actually filters first') is present, but explicit validation checkpoints before proceeding between steps are only implicit in places.

4 / 5

Progressive Disclosure

Well-organized into clear sections with one clearly signaled one-level reference to qualitative-feedback.md (and a cross-skill wikilink), and no bundle files exist locally to split further; minor gap is that the referenced file lives in a sibling skill rather than this skill's own bundle.

4 / 5

Total

15

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific, complete, and uses third-person voice with an explicit 'Use when' trigger clause. It is slightly above the midpoint on trigger-term quality only because a few natural synonyms (recordings, A/B test) are absent.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'Analyze session replay patterns across experiment variants', 'identify usability issues', 'compare behavior patterns between control and test groups', 'get qualitative insights to complement quantitative experiment results', and pairing with a linked survey — giving comprehensive coverage rather than a single generic verb.

5 / 5

Completeness

Explicitly answers 'what' ('Analyze session replay patterns across experiment variants to understand user behavior differences') and 'when' with a concrete 'Use when the user wants to...' clause enumerating several trigger situations.

5 / 5

Trigger Term Quality

Good natural-term coverage including 'session replay', 'experiment variants', 'control and test groups', 'usability issues', and 'qualitative insights', but missing common synonyms users might say such as 'session recordings', 'A/B test', or 'recordings'.

4 / 5

Distinctiveness Conflict Risk

A clear niche — analyzing session replays scoped to experiment variants via control/test groups — with distinct triggers; the survey-pairing addition stays tied to the experiment context, keeping conflict risk minimal.

5 / 5

Total

19

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 1 suspicious

Warning

referenced_paths_exist

Referenced path issues: 1 missing

Warning

Total

14

/

16

Passed

Repository
PostHog/posthog
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.