CtrlK
BlogDocsLog inGet started
Tessl Logo

exploring-ai-failures

Find where an AI/LLM application is failing in production and surface the failure patterns, working from real traces. Use when someone wants to understand what's going wrong with an AI feature, find and categorize failure modes, triage errors, or investigate quality issues (wrong answers, ignored instructions, hallucinations, tool misuse) — "what's failing in my agent", "surface error patterns", "why are the responses bad", "find the common failure modes", "what should I fix next". Covers scoping to one use case, finding failing traces by whichever signal fits the context (code errors, metric outliers, trace-type slices, manual review, existing-eval spikes, clustering), and reading them into a ranked failure taxonomy.

71

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

—

The risk profile of this skill

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A high-quality, opinionated, actionable skill body with a clear workflow and concrete tool guidance. Its main defect is broken progressive disclosure: it repeatedly offloads detail to a references/finding-traces.md file that is not present in the bundle.

Suggestions

Add the missing `references/finding-traces.md` file (referenced in Step 1 and Step 2 for taxonomy discovery and per-strategy queries) so the deferrals resolve — or inline the key queries if no bundle is intended.

Verify the `exploring-llm-traces/references/events-and-properties.md` path is correct/available, since the body depends on it for the `$ai_*` event schema.

Tighten the repeated 'loud minority / don't GROUP BY' warning so it appears fully once (e.g. Step 2's trap) and is only briefly echoed later, to recover a conciseness point.

DimensionReasoningScore

Conciseness

Mostly lean and efficient — it assumes Claude knows what an LLM/trace is and avoids concept explanation — but the 'loud minority vs silent failures / don't GROUP BY' message is restated across the intro, the Step 2 trap, Step 3, Step 4, and Tips, so a couple of those repeats could be trimmed.

4 / 5

Actionability

Gives concrete executable guidance — specific tool names with arguments (`query-llm-trace` with `traceId`, `generate-app-url {url, params}`), tool pairings (`llma-evaluation-list` + `execute-sql`), URL templates, and a 20–30 trace batch size — but the detailed query SQL is deferred to references/finding-traces.md rather than shown inline.

4 / 5

Workflow Clarity

Clear four-step sequence (Scope → Pick traces → Read batch → Rank/link/hand back) with an explicit convergence checkpoint ('Keep reading until new traces stop turning up new modes... stop when it goes quiet') and a user feedback loop ('ask which mode they want to focus on next').

5 / 5

Progressive Disclosure

The body is well-sectioned and signals references with markdown links, but the referenced `references/finding-traces.md` (cited twice for the detailed queries) and `exploring-llm-traces/references/events-and-properties.md` do not exist — no references/ directory or bundle files are present, so the promised one-level-deep navigation dead-ends.

3 / 5

Total

16

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, well-structured description: third-person voice, concrete actions, explicit 'Use when' guidance with natural quoted trigger phrases, and a distinct niche. It answers what, when, and how-to-trigger comprehensively.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'scoping to one use case', 'finding failing traces by whichever signal fits' (code errors, metric outliers, trace-type slices, manual review, existing-eval spikes, clustering), and 'reading them into a ranked failure taxonomy' — comprehensive coverage of the finding strategies.

5 / 5

Completeness

Explicitly answers both: what it does ('Find where an AI/LLM application is failing in production and surface the failure patterns, working from real traces') and when to use it ('Use when someone wants to understand what's going wrong...'), with concrete trigger phrases.

5 / 5

Trigger Term Quality

Includes natural quoted phrases a user would actually say — "what's failing in my agent", "surface error patterns", "why are the responses bad", "find the common failure modes", "what should I fix next" — alongside concrete failure types (wrong answers, hallucinations, tool misuse).

5 / 5

Distinctiveness Conflict Risk

Clear niche — production failure-pattern discovery from real traces producing a 'ranked failure taxonomy' — with distinct triggers that separate it from the related trace-reading and clustering skills.

5 / 5

Total

20

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 2 missing

Warning

referenced_paths_exist

Referenced path issues: 4 missing

Warning

Total

14

/

16

Passed

Repository
PostHog/posthog
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.