CtrlK
BlogDocsLog inGet started
Tessl Logo

exploring-llm-traces

Debug and inspect LLM/AI agent traces using PostHog's MCP tools. Use when the user pastes a trace or session URL (e.g. /ai-observability/traces/<id> or /ai-observability/sessions/<id>), asks to debug a trace, figure out what went wrong, check if an agent used a tool correctly, verify context/files were surfaced, inspect subagent behavior, investigate LLM decisions, or analyze token usage and costs. Also use when raw SQL/HogQL against `events.properties.$ai_input` / `$ai_output_choices` returns empty — message content lives only on the dedicated `posthog.ai_events` table.

68

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

—

The risk profile of this skill

SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-sequenced body with concrete executable examples and a sensible debug workflow including a truncation checkpoint. Its main weaknesses are duplicated script/date-range content and a progressive-disclosure failure: the skill references bundle files (references/, scripts/) that are not actually present.

Suggestions

Ship the referenced bundle — create references/events-and-properties.md, references/example-llm-trace.md, references/example-llm-traces-list.md, and the scripts/ files (print_summary.py, print_timeline.py, extract_span.py, extract_conversation.py, search_traces.py, show_structure.py) — so the in-body links resolve.

De-duplicate the parsing scripts: keep either the Step 4 bash walkthrough or the "Available scripts" table, not both; cross-reference one from the other.

Consolidate the date_from/date_to/timestamp anchoring guidance into Step 1 and have Step 2 reference it instead of restating the same logic.

DimensionReasoningScore

Conciseness

Mostly efficient and assumes Claude's competence, but the parsing scripts are documented twice (bash examples in Step 4 and again in the "Available scripts" table) and the date_from/date_to/timestamp anchoring logic is explained in Step 1 and re-stated in Step 2.

3 / 5

Actionability

Copy-paste-ready JSON tool-call bodies, specific bash commands with env vars (e.g. `SPAN="tool_name" python3 scripts/extract_span.py FILE`), and concrete property/table names throughout cover the common investigation cases.

5 / 5

Workflow Clarity

Clear 4-step sequence (classify URL → browse summaries → read full content → parse large results) with an explicit truncation checkpoint ("Check truncation markers before drawing conclusions ... If full detail is still truncated, narrow the query"); held below 5 because there is no full validate→fix→retry loop.

4 / 5

Progressive Disclosure

Sectioning and one-level-deep reference signaling are good, but the actual bundle is empty — references/, scripts/, and assets/ do not exist, so links like ./references/events-and-properties.md and ./scripts/print_summary.py are broken; additionally ~330 lines inline a JSON-structure tree and scripts table that belong in references.

3 / 5

Total

15

/

20

Passed

Description

95%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, third-person description that explicitly covers what the skill does and when to use it, with comprehensive natural trigger phrases and a distinctive PostHog-specific niche. The only weakness is that the core capability verbs are limited to debug/inspect, with the richer action detail living in the when-clause rather than the what-statement.

DimensionReasoningScore

Specificity

"Debug and inspect LLM/AI agent traces using PostHog's MCP tools" names the domain and core verbs, and the Use-when clause enumerates concrete subtasks ("check if an agent used a tool correctly", "verify context/files were surfaced", "analyze token usage and costs"); held below 5 because the core what-verbs are only debug/inspect.

4 / 5

Completeness

Clearly answers both what ("Debug and inspect LLM/AI agent traces using PostHog's MCP tools") and when (explicit "Use when..." with concrete trigger phrases and URL examples).

5 / 5

Trigger Term Quality

Rich natural triggers users would actually say — "pastes a trace or session URL", "asks to debug a trace", "figure out what went wrong", "inspect subagent behavior", "analyze token usage and costs", plus the SQL/HogGL empty-result case — with synonyms and concrete URL paths.

5 / 5

Distinctiveness Conflict Risk

"PostHog's MCP tools" combined with specific URL paths (/ai-observability/traces/<id>) and the posthog.ai_events table carves a clear niche with minimal conflict risk against other skills.

5 / 5

Total

19

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 11 missing

Warning

referenced_paths_exist

Referenced path issues: 12 missing

Warning

Total

14

/

16

Passed

Repository
PostHog/posthog
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.