CtrlK
BlogDocsLog inGet started
Tessl Logo

exploring-ai-failures

Find where an AI/LLM application is failing in production and surface the failure patterns, working from real traces. Use when someone wants to understand what's going wrong with an AI feature, find and categorize failure modes, triage errors, or investigate quality issues (wrong answers, ignored instructions, hallucinations, tool misuse) — "what's failing in my agent", "surface error patterns", "why are the responses bad", "find the common failure modes", "what should I fix next". Covers scoping to one use case, finding failing traces by whichever signal fits the context (code errors, metric outliers, trace-type slices, manual review, existing-eval spikes, clustering), and reading them into a ranked failure taxonomy.

70

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-organized, actionable skill body with a clear four-step workflow and proper progressive disclosure into a real reference file. Minor redundancy in the Tips section and conceptual-level treatment of some strategies are the only drags on an otherwise strong skill.

Suggestions

Trim the Tips section to points not already covered in Steps 1–4 to reduce redundancy and earn back conciseness tokens.

Inline one concrete example query per finding strategy (or a one-line pointer to the exact reference section) so each strategy is immediately executable from the body.

Make the Step 3 stop criterion and Step 4 verification loop explicit as labeled checkpoints (e.g., 'Checkpoint: new modes have stopped emerging') to strengthen workflow clarity.

DimensionReasoningScore

Conciseness

Lean and opinionated, assuming Claude's competence, but the Tips section largely restates points already made in the steps (reading is the job, don't over-index on errors, strategies are a menu), adding mild redundancy.

4 / 5

Actionability

Provides concrete, executable guidance — a tools table, exact generate-app-url call templates with params, and the traceId argument — but several finding strategies are described at a conceptual level with the executable queries deferred to the reference file.

4 / 5

Workflow Clarity

A clearly sequenced four-step process (scope → pick → read → rank) with a stop criterion ('keep reading until new traces stop turning up new modes') and a self-correction loop (a flagged hallucination may be correct in context); minor checkpoint explicitness gaps keep it from 5.

4 / 5

Progressive Disclosure

Well-structured overview with a clearly signaled one-level-deep reference (references/finding-traces.md, which exists) holding the detailed queries, plus organized sections; the fairly long inline body keeps it just short of a 5.

4 / 5

Total

16

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, concrete description with explicit what/when guidance and a rich set of natural trigger phrases in third-person voice. Only minor overlap risk with closely related trace-reading skills keeps distinctiveness from a 5.

DimensionReasoningScore

Specificity

Names multiple concrete actions — 'Find where an AI/LLM application is failing', 'surface the failure patterns', 'scoping to one use case', 'finding failing traces', 'reading them into a ranked failure taxonomy' — with comprehensive coverage of the workflow.

5 / 5

Completeness

Explicitly answers both what (find and surface failure patterns, build a ranked failure taxonomy) and when ('Use when someone wants to understand what's going wrong with an AI feature...') with concrete trigger phrases.

5 / 5

Trigger Term Quality

Includes natural user phrases verbatim — "what's failing in my agent", "surface error patterns", "why are the responses bad", "find the common failure modes", "what should I fix next" — plus synonyms like 'triage errors' and 'investigate quality issues', giving comprehensive natural-term coverage.

5 / 5

Distinctiveness Conflict Risk

Carves a clear niche around production failure investigation and ranked failure modes, but overlaps slightly with sibling trace skills (exploring-llm-traces, exploring-llm-clusters) that share the trace-reading substrate.

4 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
PostHog/posthog
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.