CtrlK
BlogDocsLog inGet started
Tessl Logo

diagnosing-bugs

Diagnosis loop for hard bugs and performance regressions. Use when the user says "diagnose"/"debug this", or reports something broken/throwing/failing/slow.

74

Quality

91%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

High

Do not use without reviewing

The canonical home for this skill is diagnosing-bugs in coder/agent-tty

SKILL.md
Quality
Evals
Security

Quality

Content

96%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exceptionally well-crafted process skill: lean, imperative prose with zero concept padding, hard phase gates with explicit completion criteria and checklists, and concrete commands and formats at every step. The only real gap is progressive disclosure — the long loop-construction recipe list is inlined where a reference file would keep the core tighter.

DimensionReasoningScore

Conciseness

The body never explains concepts Claude already knows (git bisect, Playwright, RNG seeding, HAR files are all name-checked without tutorial text) and is relentlessly directive — 'Tag every debug log with a unique prefix… Untagged logs survive; tagged logs die.' Every section changes behavior rather than padding context. Not 4 because the few rhetorical moments ('Be aggressive. Be creative. Refuse to give up.') are emphasis that shapes behavior, not over-explanation that could be trimmed.

5 / 5

Actionability

Concrete, executable guidance throughout: 'Tag every debug log with a unique prefix, e.g. [DEBUG-a4f2]', 'a single function call', 'git bisect run it', 'run 1000 random inputs', 'paste the invocation and its output', and the falsifiable hypothesis format 'If <X> is the cause, then <changing Y> will make the bug disappear'. Per the scoring notes, an instruction-only skill with this density of specific commands and formats fully qualifies. Not 4 because there are no pseudocode gaps — every instruction names the tool, format, or command to use.

5 / 5

Workflow Clarity

Six clearly sequenced phases, each with an explicit completion criterion and hard gates between them: 'No red-capable command, no Phase 2', 'Do not proceed until you have reproduced and minimised', and a required pre-done checklist in Phase 6 ('grep the prefix', 're-run the Phase 1 loop'). This matches the top anchor — explicit validation steps, feedback loops, and checklists. Not 4 because checkpoints are not merely present; they block progression and re-verify earlier phases (Phase 5 step 5 re-runs the Phase 1 loop).

5 / 5

Progressive Disclosure

Good structure: clear phase headers, one-level-deep reference to scripts/hitl-loop.template.sh (verified to exist in the bundle), referenced at exactly the point of use. However, at 134 lines the body inlines the 10-item 'Ways to construct one' recipe list — reference-style material that could live in a references/ file so the core SKILL.md stays a lean process overview. Not 5 because content that could be split out is inline; not 3 because what is inline is well-organized and the one reference that exists is clearly signaled.

4 / 5

Total

19

/

20

Passed

Description

86%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: terse, third-person, with an explicit and unusually well-worded 'Use when' clause containing natural trigger synonyms. Its only weakness is specificity — it names the process ('diagnosis loop') rather than the concrete actions inside it — plus slight over-breadth in the trigger terms.

Suggestions

List 2-3 of the skill's concrete actions in the description (e.g. 'reproduce and minimise, generate ranked hypotheses, instrument, then fix with a regression test') to lift the specificity beyond 'Diagnosis loop'.

Tighten the trigger scope so ordinary bug-fix requests don't fire it — e.g. add 'hard'/'stubborn'/'regression' qualifiers to the reported-symptom terms.

Consider mentioning the performance-regression branch earlier with a natural term like 'slow' already present — currently 'slow' is the only perf trigger and sits at the end of a long list.

DimensionReasoningScore

Specificity

Names the domain ('hard bugs and performance regressions') and one action ('Diagnosis loop') but does not list the several concrete actions the skill performs (reproduce, minimise, hypothesise, instrument, fix + regression test), so it matches the 1-2-concrete-actions anchor rather than the several-specific-actions anchor above.

3 / 5

Completeness

Explicitly answers both what ('Diagnosis loop for hard bugs and performance regressions') and when ('Use when the user says "diagnose"/"debug this", or reports something broken/throwing/failing/slow') with concrete trigger phrases — a direct match for the top anchor.

5 / 5

Trigger Term Quality

'diagnose', 'debug this', 'broken', 'throwing', 'failing', 'slow' are exactly the natural phrases and synonyms users say when reporting a bug, giving comprehensive coverage of the domain's natural vocabulary. Not 4 because no common variation of a bug/perf-regression report is missing.

5 / 5

Distinctiveness Conflict Risk

The 'hard bugs and performance regressions' scope carves a niche, but 'broken/throwing/failing/slow' would fire on almost any bug report, creating minor overlap with general debugging and fix-it workflows. Not 5 because the trigger surface is broader than a clearly distinct niche.

4 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
udecode/plate
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.