CtrlK
BlogDocsLog inGet started
Tessl Logo

debugging

Runs a hypothesis-driven debugging loop across any language or binary, escalating to orthogonal oracle angles and locking the fix with a failing test. Use for crashes, silent failures, hangs, wrong responses, memory leaks, async misbehavior, or reverse engineering.

66

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an exemplary progressive-disclosure index with a well-sequenced, validation-gated phase workflow and verified, one-level-deep references. Its one real weakness is redundancy: the mandate to read the references is repeated in roughly six blocks across the file, which inflates token cost without adding information.

Suggestions

Consolidate the repeated read-the-references enforcement (the 🚨 headline block, the intro blockquote, the gate rule, and the re-statements in the Runtime Setup and Phase Loop sections) into a single stated gate rule; the message currently appears ~6 times.

Trim table cells that re-explain the enforcement itself (e.g., "ACTIVELY USE WHEN THE SCENARIO FITS" and "Failing to use these tools in their domain is a process failure") down to one statement, keeping the per-tool gotcha payloads that earn their tokens.

DimensionReasoningScore

Conciseness

The index design teaches nothing Claude already knows and the tables are dense with non-obvious gotcha facts ("tsx + node inspect CLI has a silent source-map failure", "strings -n 8 silently drops short interpolations"), but the read-the-references enforcement message is repeated in ~6 separate blocks (the 🚨 headline, the intro blockquote, the gate rule, the runtime-setup section, the phase-loop section, and "What to Do Right Now"). This fits anchor 3 — mostly efficient but with repeated exhortation that could be consolidated into one stated gate rule. It is above anchor 2 because the repeated content is a deliberate behavioral pressure, not explanation of known concepts, and the tables carry real payload.

3 / 5

Actionability

Concrete decision tables map runtime → reference, tool → reference, and phase → reference with exact file paths; the native-vs-bundled discriminator is fully executable ("du -h ./target" plus "strings -n 12 ./target | rg -iE 'bun|node_modules|...'"); and "What to Do Right Now" is a numbered executable entry sequence. It stops short of anchor 5 because the bulk of copy-paste commands deliberately lives in the references and some directives are behavioral rather than executable ("spawn three Oracles with orthogonal framings").

4 / 5

Workflow Clarity

The phase loop gives a clearly sequenced 0-10 workflow with each step mapped to its reference, and validation checkpoints are explicit: Phase 7 requires a proof "that fails on the pre-fix code", Phase 9 requires verifying "git diff shows only fix + test", and Phase 10 names evidence gates before declaring done. Feedback loops are present (Phase 4 escalates to the Oracle Triple after two consecutive failed rounds), matching anchor 5.

5 / 5

Progressive Disclosure

This is a model index-over-knowledge design: the body states "The knowledge is in references/" and every one of the 22 referenced files exists in the bundle (verified by listing), each link is one level deep with a clear when-to-read condition, and navigation is by decision table rather than prose. The deep link to "partial-runtime-evidence.md#verification-oracle-pattern-for-non-debug-tasks" resolves to a real heading in that file. Fully matches anchor 5.

5 / 5

Total

17

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with an explicit what-and-when structure and natural trigger terms covering the main debugging failure modes. Its weaknesses are the unexplained internal jargon "orthogonal oracle angles" and the absence of the most common user-facing trigger words (bug, error, stack trace).

DimensionReasoningScore

Specificity

Names three concrete actions spanning the debugging lifecycle — "Runs a hypothesis-driven debugging loop", "escalating to orthogonal oracle angles", "locking the fix with a failing test" — with only a minor gap: "orthogonal oracle angles" is internal jargon a reader outside the skill cannot decode, and verification/cleanup steps go unmentioned. It sits below anchor 5 (fully comprehensive, all actions concrete) but above anchor 3 because more than 1-2 specific actions are listed.

4 / 5

Completeness

Explicitly answers both halves: the "what" is the concrete three-part debugging loop, and "Use for crashes, silent failures, hangs, wrong responses, memory leaks, async misbehavior, or reverse engineering" is a direct when-to-use clause with concrete trigger phrases, matching the anchor-5 example structure exactly. Not below 4 because neither half is missing or implicit.

5 / 5

Trigger Term Quality

"crashes, silent failures, hangs, wrong responses, memory leaks, async misbehavior, or reverse engineering" gives strong natural phrases users actually say when a debugging skill is needed. It falls short of anchor 5's comprehensive synonym coverage: the most common user words "bug", "debug" (as a user request), "error", "stack trace", and "exception" are absent.

4 / 5

Distinctiveness Conflict Risk

Malfunction-mode triggers (crashes, hangs, memory leaks, silent failures) form a clear niche distinct from review or build skills, but "reverse engineering" and "wrong responses" create minor overlap risk with security-review and LLM-behavior skills. It is more distinct than anchor 3 ("could still overlap with similar skills") but not the minimal-conflict clarity of anchor 5.

4 / 5

Total

17

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 26 deeper-than-1-level

Warning

referenced_paths_exist

Referenced path issues: 53 deeper-than-1-level

Warning

Total

14

/

16

Passed

Repository
code-yeongyu/lazycodex
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.