CtrlK
BlogDocsLog inGet started
Tessl Logo

agent-introspection-debugging

Structured self-debugging workflow for AI agent failures using capture, diagnosis, contained recovery, and introspection reports. Use when an agent run fails and you need a reproducible diagnosis instead of a retry.

74

Quality

91%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

96%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, highly actionable workflow skill: concrete templates, a specific failure-pattern table, and an explicitly validated four-phase sequence. Its only real costs are redundancy between the recovery heuristics and the phase sections, and a monolithic single-file layout that could shed the templates to a reference file.

DimensionReasoningScore

Conciseness

The body is efficient and free of concept explanations Claude already knows, but the 'Recovery Heuristics' section substantially restates Phase 2/3 guidance ('Verify the world state instead of trusting memory' duplicates 're-check the actual filesystem / branch / process state'; 'Shrink the failing scope' duplicates 'narrow the task to one failing command'), and the good/bad pattern section repeats it a third time.

4 / 5

Actionability

Three copy-paste-ready markdown templates (failure capture, recovery action, self-debug report), a diagnosis table with concrete failure signatures (ECONNREFUSED, 429, file missing after write, tests still failing after 'fix'), and specific diagnosis questions make the guidance fully executable for an instruction-only skill.

5 / 5

Workflow Clarity

The four-phase loop (Capture -> Diagnose -> Recover -> Report) is explicitly sequenced with validation checkpoints ('What evidence would prove the fix worked', 'Run one discriminating check') and feedback loops ('change the plan only if the check supports it', 'Only then retry'), plus a defined output standard.

5 / 5

Progressive Disclosure

A single well-organized file with clear section headers and easy navigation, but at ~150 lines everything is inline; the pattern table and the three templates are sizable enough that a one-level-deep references/ split would fit the level-5 anchor, and no bundle files exist to offload them.

4 / 5

Total

18

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that names the domain, enumerates concrete capabilities, and gives an explicit use-when trigger. The only deductions are the second-person phrasing in the trigger clause and a few missing natural synonyms for failure states.

Suggestions

Rewrite the trigger clause in third person to avoid the specificity penalty, e.g. 'Use when an agent run fails and a reproducible diagnosis is needed instead of another retry.'

Add natural trigger synonyms users would actually say, such as 'stuck in a loop', 'retrying repeatedly', or 'burning tokens without progress', to broaden keyword coverage.

DimensionReasoningScore

Specificity

The description lists four concrete actions covering the full workflow ('capture, diagnosis, contained recovery, and introspection reports'), which matches the level-5 anchor, but the trigger clause 'you need a reproducible diagnosis instead of a retry' is second person, incurring the rubric's 1-point voice penalty on specificity.

4 / 5

Completeness

Explicitly answers both: what ('Structured self-debugging workflow for AI agent failures using capture, diagnosis, contained recovery, and introspection reports') and when ('Use when an agent run fails and you need a reproducible diagnosis instead of a retry'), with concrete trigger phrases — the exact level-5 anchor shape.

5 / 5

Trigger Term Quality

Includes natural phrases users would say ('agent run fails', 'retry', 'self-debugging', 'diagnosis'), matching 'good keyword coverage; a few natural terms missing'. It lacks common variations such as 'stuck', 'looping', 'runaway', or 'token burn', so it falls short of the comprehensive-synonym level 5.

4 / 5

Distinctiveness Conflict Risk

'AI agent failures' plus 'self-debugging workflow' carves a clear niche with distinct triggers, and it explicitly positions against retrying, so it is unlikely to fire for general code debugging or verification skills.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
affaan-m/ECC
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.