CtrlK
BlogDocsLog inGet started
Tessl Logo

systematic-debugging

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes

49

Quality

53%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

High

Do not use without reviewing

Fix and improve this skill with Tessl

tessl review fix ./eval/local/skills/benchmarks/dependency/superpowers/systematic-debugging/SKILL.md

The canonical home for this skill is systematic-debugging in obra/superpowers

SKILL.md
Quality
Evals
Security

Quality

Content

58%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content delivers an excellent, well-validated debugging workflow with genuinely actionable instrumentation examples, but it is padded with repetitive motivational material and unsupported statistics, and all three of its supporting-technique file references are dangling because those files are absent from the bundle. Trimming the exhortation sections and shipping (or removing) the referenced files would fix the two weakest dimensions.

Suggestions

Add the missing bundle files `root-cause-tracing.md`, `defense-in-depth.md`, and `condition-based-waiting.md` (or remove the references) — every referenced path currently dead-ends.

Cut the repetitive exhortation: consolidate "Red Flags", "your human partner's Signals", and "Common Rationalizations" into a single short section, and delete the unsupported "Real-World Impact" statistics.

Move the long multi-layer signing example and the rationalization tables into a separate reference file so SKILL.md stays a lean overview.

DimensionReasoningScore

Conciseness

The body repeats the same exhortation many times ("STOP. Return to Phase 1" appears in the Red Flags, Signals, and multiple phase sections) and includes padded sections like "Common Rationalizations", "Real-World Impact" with unverifiable statistics ("First-time fix rate: 95% vs 40%"), and the cryptic line "Violating the letter of this process is violating the spirit of debugging". This is noticeably verbose with several unnecessary sections (anchor 2) rather than only some tightenable spots (anchor 3).

2 / 5

Actionability

Guidance is mostly executable: a concrete, runnable bash instrumentation example (env inspection, `security list-keychains`, `codesign --sign "$IDENTITY" --verbose=4`), specific commands like `git diff`, and explicit numbered steps per phase. The Phase 1.4 instrumentation block is structured pseudocode rather than executable code, but that is justified since it is generic to any multi-component system; minor gaps keep it at anchor 4 rather than 5.

4 / 5

Workflow Clarity

The four phases are strictly sequenced ("You MUST complete each phase before proceeding to the next") with explicit validation checkpoints and feedback loops: hypothesis verification ("Did it work? Yes → Phase 4 / Didn't work → form NEW hypothesis"), fix verification ("Test passes now? No other tests broken?"), an escalation rule after 3 failed fixes, and a Quick Reference table. This matches anchor 5's clear sequence with explicit validation and error-recovery loops.

5 / 5

Progressive Disclosure

References to `root-cause-tracing.md`, `defense-in-depth.md`, and `condition-based-waiting.md` are clearly signaled and one level deep, but none of these files exist in the bundle (no references/, scripts/, or assets/ directories are present), so navigation dead-ends. Combined with substantial inline content (rationalization tables, multi-layer bash example, motivational sections) that belongs in separate files, this sits at anchor 2 rather than 3, where references at least resolve.

2 / 5

Total

13

/

20

Passed

Description

47%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description has strong, explicit trigger guidance with natural user phrasing, but it completely omits any statement of what the skill does (e.g., that it enforces a systematic four-phase root-cause investigation process). Adding a capability clause before the "Use when" clause would raise both completeness and specificity substantially.

Suggestions

Lead with a concrete capability statement, e.g., "Systematically debug by investigating root causes through a four-phase process (investigation, pattern analysis, hypothesis testing, implementation). Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes."

Add common trigger synonyms such as "error", "crash", "stack trace", "debugging", or "tests are failing" to broaden natural keyword coverage.

Mentioning the signature constraint ("find root cause before attempting any fix") in the description would sharpen distinctiveness against generic troubleshooting skills.

DimensionReasoningScore

Specificity

The description names the debugging domain via concrete scenarios ("any bug, test failure, or unexpected behavior") but states no capability or action whatsoever — there is no verb describing what the skill does. It falls between anchor 1 (pure abstract language) and anchor 2 (domain named, actions minimal/generic), and closer to 2 because the trigger scenarios are concrete rather than vague fluff.

2 / 5

Completeness

The "when" is explicit and well-formed ("Use when encountering any bug, test failure, or unexpected behavior"), but the "what" is entirely absent — the description never says what the skill does. This matches anchor 2 ("only 'when' is present without 'what'") and cannot be 3, which requires a clear 'what'.

2 / 5

Trigger Term Quality

"bug", "test failure", and "unexpected behavior" are natural phrases users would say when they need this skill, giving good keyword coverage. Not 5 because common synonyms and variations like "error", "crash", "debugging", or "not working" are missing; not 3 because the terms present go beyond partial coverage.

4 / 5

Distinctiveness Conflict Risk

Bug, test failure, and unexpected behavior are clearly debugging-niche triggers, making it mostly distinct from unrelated skills. Not 5 because "unexpected behavior" is broad enough to overlap with other troubleshooting, testing, or code-review skills.

4 / 5

Total

12

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
rpamis/comet
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.