CtrlK
BlogDocsLog inGet started
Tessl Logo

systematic-debugging

Investigates a test failure, bug, or unexpected behavior to find its root cause before proposing any fix - reproduce the issue, trace the data flow, form and test one hypothesis at a time, then implement and verify the fix. This is the debugging stage inside the delivery-flow workflow, entered by requests like "debug this failure", "why is this test failing", "find the root cause", or "this bug keeps coming back" once delivery-flow has handed off a bug. Do not activate directly for a standalone bug report that has not gone through delivery-flow first.

68

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

The canonical home for this skill is tessleng/sdlc-implementation

SKILL.md
Quality
Evals
Security

Quality

Content

73%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers an exceptionally clear four-phase debugging workflow with strong validation checkpoints, feedback loops, and real one-level-deep reference files. Its main weakness is token efficiency: roughly a third of the content repetitively enforces the same "root cause first" message across three overlapping sections, and two reference files lack in-context pointers.

Suggestions

Collapse the "Red Flags", "Common Rationalizations", and "your human partner's Signals" sections into a single compact table — they all restate the same no-fix-without-root-cause enforcement, which would cut ~60 lines without losing guidance.

Add in-context pointers to defense-in-depth.md and condition-based-waiting.md at the phases where they apply (e.g., after Phase 4 fix verification, and during waiting/retry scenarios) rather than only listing them at the end.

Move the domain-specific multi-layer codesign example into a reference file (e.g., references/instrumentation-example.md) and keep only the generic boundary-instrumentation checklist inline.

DimensionReasoningScore

Conciseness

The four-phase core is dense and instructional, but the body pads the same enforcement message across three redundant sections — "Red Flags", "Common Rationalizations", and "your human partner's Signals" all restate "no fixes without root cause" — plus filler like "Violating the letter of this process is violating the spirit of debugging." This fits the mostly-efficient-but-could-be-tightened anchor; it is not the severely-verbose anchor at 1 because nothing explains concepts Claude already knows, and not 4 because the redundancy is substantial, not minor.

3 / 5

Actionability

Concrete guidance includes a copy-paste bash instrumentation example for multi-component systems, an explicit numeric decision rule ("If ≥ 3: STOP and question the architecture"), a failing-test gate, and pointers to real scripts like find-polluter.sh. It stops short of 5 because some directives remain abstract ("Trace Data Flow", "Find Working Examples") without concrete techniques for the non-referenced phases, and the inline example is domain-specific (macOS codesign) rather than covering common cases.

4 / 5

Workflow Clarity

The four phases are strictly sequenced with "You MUST complete each phase before proceeding to the next", each phase has numbered sub-steps, a Quick Reference table gives success criteria per phase, and there is an explicit feedback loop: verify fix → if failed and <3 attempts return to Phase 1 → if ≥3 attempts stop and question architecture with the human partner. This matches the clear-sequence-with-explicit-validation-and-feedback-loops anchor.

5 / 5

Progressive Disclosure

All three referenced files (root-cause-tracing.md, defense-in-depth.md, condition-based-waiting.md) exist, are one level deep, and are clearly listed with one-line descriptions in Supporting Techniques; root-cause-tracing is also signaled at point of use in Phase 1. It falls short of 5 because defense-in-depth and condition-based-waiting have no in-context pointers where they would apply, and the long inline codesign instrumentation example could itself live in a reference file.

4 / 5

Total

16

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: specific, third-person, with explicit trigger phrases and an unusually clear activation boundary tied to the delivery-flow workflow. Its only weakness is slightly narrow trigger coverage that omits common synonyms such as "regression" or "crash".

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — "reproduce the issue, trace the data flow, form and test one hypothesis at a time, then implement and verify the fix" — covering the full investigation-to-fix lifecycle comprehensively. It is not merely naming the debugging domain; every stage is a specific instruction, matching the comprehensive-coverage anchor rather than the several-actions-with-minor-gaps anchor below.

5 / 5

Completeness

It explicitly answers what the skill does ("Investigates a test failure... to find its root cause before proposing any fix") and when to use it, with concrete trigger phrases and an explicit entry condition ("once delivery-flow has handed off a bug"). Both halves are present with concrete trigger phrases, matching the top anchor exactly.

5 / 5

Trigger Term Quality

It quotes natural phrases users would actually say — "debug this failure", "why is this test failing", "find the root cause", "this bug keeps coming back" — giving good keyword coverage. It falls short of the 5 anchor because common synonyms like "regression", "crash", or "this is broken" are missing, though it clearly exceeds the some-keywords-missing-variants anchor at 3.

4 / 5

Distinctiveness Conflict Risk

It carves a clear niche as "the debugging stage inside the delivery-flow workflow" and adds an explicit negative boundary — "Do not activate directly for a standalone bug report that has not gone through delivery-flow first" — which minimizes conflict risk with generic debugging skills. This is distinct beyond the minor-overlap anchor at 4.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
tesslio/tessl-eval-demo
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.