CtrlK
BlogDocsLog inGet started
Tessl Logo

systematic-debugging

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes

52

Quality

56%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./plugins/superpowers/skills/systematic-debugging/SKILL.md

The canonical home for this skill is systematic-debugging in obra/superpowers

SKILL.md
Quality
Evals
Security

Quality

Content

70%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers an excellent, well-gated workflow — clear phase sequencing, explicit validation, and strong feedback loops — with largely concrete, actionable guidance. Its weaknesses are redundancy (the same stop-and-reinvestigate message repeated across four sections plus a summary table) and progressive disclosure: it points to three supporting technique files that do not exist in the bundle.

Suggestions

Create the referenced bundle files (root-cause-tracing.md, defense-in-depth.md, condition-based-waiting.md) or remove the references — currently the "Supporting Techniques" section and the Phase 1 pointer navigate to files that are not present.

Consolidate the "Red Flags", "Common Rationalizations", and "your human partner's Signals" sections into a single anti-pattern section (or move them to a reference file); they repeat the same "return to Phase 1" message three times.

Replace the pseudocode instrumentation block ("For EACH component boundary: Log what data enters...") with a concrete runnable command template, mirroring the style of the bash example that follows it.

DimensionReasoningScore

Conciseness

The body never explains concepts Claude already knows, but the same message ("STOP, return to Phase 1, don't guess") is repeated across four sections — "Don't skip when", "Red Flags", "your human partner's Signals", and the "Common Rationalizations" table — and the Quick Reference table restates the four phases. This is anchor 3 ('mostly efficient but could be tightened') rather than 4, given the deliberate multi-section redundancy.

3 / 5

Actionability

Guidance is concrete and specific: an executable bash instrumentation example (env | grep IDENTITY, security find-identity -v, codesign --verbose=4), explicit git-diff checks, and precise heuristics ("if >= 3 fixes: question the architecture"). The one gap is that the component-boundary instrumentation block is pseudocode ("For EACH component boundary: Log what data enters..."), which fits anchor 4 ('mostly executable guidance, minor gaps') rather than 5.

4 / 5

Workflow Clarity

The four phases are explicitly sequenced with a hard gate ("You MUST complete each phase before proceeding"), validation checkpoints ("Test passes now? No other tests broken?"), and explicit feedback loops (failed fix -> return to Phase 1; >= 3 failures -> architecture review with human-partner escalation). This matches anchor 5 exactly: clear sequence, explicit validation, feedback loops, and a summary checklist.

5 / 5

Progressive Disclosure

Section structure is good and the supporting techniques are clearly signaled ("See root-cause-tracing.md in this directory"), but all three referenced files (root-cause-tracing.md, defense-in-depth.md, condition-based-waiting.md) are missing from the directory, leaving broken navigation. Combined with ~280 lines of inline red-flag/rationalization content that could live in a separate file, this fits anchor 3 ('some structure, references present but the organization could improve') rather than 4, whose structure must actually hold together.

3 / 5

Total

15

/

20

Passed

Description

43%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description has a strong, explicit trigger clause with good natural trigger terms, but it omits the "what" entirely — a user reading it knows when to invoke the skill but not what it will do. Adding a statement of the skill's core capability (systematic root-cause investigation before fixing) would lift both completeness and specificity.

Suggestions

Add a concrete "what" clause, e.g. "Investigates root causes through a four-phase process (reproduce, analyze, hypothesize, fix) before any code changes. Use when encountering any bug..." — this would raise completeness from 2 toward 4-5.

Include common trigger synonyms such as "error", "crash", "failing tests", or "not working as expected" to broaden natural keyword coverage.

Narrow the scope qualifier from "any bug" to something like "any bug or test failure you are about to fix" to reduce overlap with related review/fixing skills.

DimensionReasoningScore

Specificity

The description names the domain ("any bug, test failure, or unexpected behavior") but describes no concrete actions — what the skill actually does (root cause investigation, the four-phase process) is entirely absent. It fits anchor 2 ('names the domain but actions are minimal or generic') rather than 3, which requires 1-2 concrete actions to be listed.

2 / 5

Completeness

Only the "when" is present — "Use when encountering any bug, test failure, or unexpected behavior" — with no statement of what the skill does; "before proposing fixes" gestures at behavior but names no capability. This matches anchor 2 ('only when is present without what') exactly; it cannot be 3 because the "what" is not merely weakly implied, it is absent.

2 / 5

Trigger Term Quality

"bug", "test failure", and "unexpected behavior" are natural phrases users would say when they need this skill. Common variations like "error", "crash", "failing tests", and "not working" are missing, so it fits anchor 4 ('good keyword coverage; a few natural terms missing') rather than 5's comprehensive synonym coverage.

4 / 5

Distinctiveness Conflict Risk

The debugging niche is somewhat distinct, but "any bug" and "unexpected behavior" are broad triggers that could overlap with code-fixing, code-review, or TDD skills. This sits at anchor 3 ('somewhat specific but could still overlap with similar skills'), below 4 because the word "any" widens the trigger surface considerably.

3 / 5

Total

11

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
openai/plugins
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.