CtrlK
BlogDocsLog inGet started
Tessl Logo

debugging-difficult-bugs

Use early when debugging a medium or hard bug, especially when tests alone may not reveal the real runtime failure. Trigger this before extended TDD iteration when a bug involves runtime state, ordering, persistence, streaming, concurrency, UI/manual reproduction, external services, or when a red or newly passing test may not model the real issue. Skip only when the root cause is already directly proven by a stack trace or deterministic test that exercises the real runtime path.

68

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, self-contained workflow skill: the required approach is sequenced with validation gates, the logging pattern is executable as written, and verification is concrete rather than hand-wavy. Weaknesses are minor — mild redundancy between the description, 'When to Use', and the reproduction sections, and a single-language (Node/TS) code example.

Suggestions

Merge the step-3 'Reproduce the real issue' guidance with the 'Reproduction Handoff to User' section — the same instructions appear twice and could be stated once.

Trim the 'When to Use' bullet list, which largely restates the frontmatter description; keep only the triggers not already in the description (e.g. the temptation-to-make-a-second-speculative-fix cues).

Consider a short one-line note that the append-only JSONL pattern adapts to other languages (Python `open(..., 'a')`, etc.) rather than only Node/TS, or trim the second code block if Node-only is intentional.

DimensionReasoningScore

Conciseness

Dense, list-driven, and free of concepts Claude already knows (no explanation of what JSONL or concurrency is). Minor over-explanation remains: the 'When to Use' section restates the frontmatter description almost verbatim, and reproduction guidance appears in both step 3 and the 'Reproduction Handoff to User' section.

4 / 5

Actionability

Copy-paste-ready `debugBug` logger implementation, concrete call-site examples covering success, error, and redaction cases, an exact message template for the user handoff, and specific multi-process file-naming guidance — fully executable with the common cases covered.

5 / 5

Workflow Clarity

Six clearly sequenced steps with explicit validation checkpoints ('Analyze the log before fixing', 'Only then implement the fix'), an analysis checklist of divergence questions, and a final verification checklist ('regression test fails before the fix and passes after', 'final diff contains only the fix and intentional tests') — full feedback loops from instrument → reproduce → analyze → fix → verify → clean up.

5 / 5

Progressive Disclosure

No bundle files exist (no references/, scripts/, or assets/), and the body references no external files — so there are no dangling references and no nested-reference risk. Sections are well-organized and all inline content is core to the workflow, but the skill is ~175 lines (above the under-50-line simple-skill exception) and the Node/TS code block is a candidate for trimming or extraction.

4 / 5

Total

18

/

20

Passed

Description

77%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An unusually strong trigger-focused description with explicit use-early and skip-only-when boundary guidance and rich natural keyword coverage. Its main weakness is that the 'what' is characterized only at the domain level — the skill's actual approach (unconditional JSONL instrumentation, real reproduction, log analysis before fixing) never appears.

Suggestions

Add one clause naming the concrete methodology, e.g. 'Instruments the real runtime path with unconditional JSONL logs, reproduces the issue, and analyzes the log before fixing' — this would lift both specificity and completeness.

Tighten the long trigger enumeration slightly; 'ordering, persistence, streaming, concurrency' can be compressed (e.g. 'runtime ordering, state, concurrency, or external services') without losing natural keywords.

DimensionReasoningScore

Specificity

The description names the domain ("debugging a medium or hard bug", "tests alone may not reveal the real runtime failure") but enumerates no concrete actions of the skill itself — instrumentation, JSONL logging, reproduction, and log analysis before fixing are all absent, matching 'names domain and 1-2 concrete actions, but not comprehensive'.

3 / 5

Completeness

The 'when' is explicit and rich with concrete trigger phrases ("Use early when…", "Trigger this before extended TDD iteration…", "Skip only when…"), but the 'what' stops at the domain level — the actual debugging methodology the skill prescribes is never stated. Not 5 because the what-half lacks concrete actions; not 3 because the when-half is explicit, not weakly implied.

4 / 5

Trigger Term Quality

Comprehensive natural-term coverage: "debugging", "red" test, "newly passing test", "reproduce" manually, "runtime state", "ordering", "persistence", "streaming", "concurrency", "external services", "stack trace", "deterministic test" — including synonyms (red test / newly passing test / failing test may not model the real issue).

5 / 5

Distinctiveness Conflict Risk

A clear niche (bugs where tests may not model the real runtime failure) with an explicit skip clause that carves it apart from generic debugging and TDD skills; minor residual overlap because "debugging a medium or hard bug" would also naturally trigger a general debugging skill.

4 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
mastra-ai/mastra
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.