CtrlK
BlogDocsLog inGet started
Tessl Logo

ce-debug

Diagnosis loop for bugs and failing behavior. Use for errors, stack traces, regressions, failed tests, issue-tracker bugs, stuck investigations after failed fixes, or asks to debug/fix a bug.

71

Quality

87%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-3

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, highly actionable debugging workflow with clear phased sequencing, validation gates, feedback loops, and clean progressive disclosure into four real reference files. The only meaningful weakness is conciseness — the prose is longer than necessary in places and restates inferences Claude could make on its own.

Suggestions

Tighten the explanatory prose in Phase 1.3 and the Core Principles: drop restatements Claude can infer (e.g. 'The bottom frame is the symptom; the root cause is somewhere upstream') and keep only the non-obvious methodology.

Condense the repeated per-platform blocking-tool name lists (Claude Code/Codex/Antigravity/Pi appear verbatim in both Phase 2 and Phase 4) into a single defined reference once, then point back to it.

Merge the long Phase 4 polish/review-tail paragraphs into a scannable checklist; much of the conditional branching reads as prose that could be bulletized to save tokens without losing clarity.

DimensionReasoningScore

Conciseness

The body is efficient for a complex multi-phase skill and is mostly domain-specific rather than explaining basics Claude lacks, but it runs ~325 lines with several prose passages that restate reasoning Claude can infer (e.g. 'The bottom frame is the symptom; the root cause is somewhere upstream') and could be tightened.

2 / 3

Actionability

Highly actionable for an instruction-only skill: concrete commands (gh issue view, git bisect, git rev-parse, git checkout -b), structured output templates (Debug Summary, Post-Fix Quality), and explicit per-platform tool names — copy-paste ready guidance throughout.

3 / 3

Workflow Clarity

Clear 0–4 phase sequence with an overview table, explicit gates (causal-chain gate, fix-choice gate, workspace check), validation checkpoints (test fails for the right reason, passes, broader suite, re-verify after tail edits), and feedback loops (failed fix returns to Phase 2 with hypothesis invalidation; 3 failed attempts escalate).

3 / 3

Progressive Disclosure

SKILL.md is an overview pointing to four real, one-level-deep reference files (pipeline-mode, investigation-techniques, anti-patterns, defense-in-depth), each conditionally loaded with clear 'Read references/...' signals; content is appropriately split and easy to navigate.

3 / 3

Total

11

/

12

Passed

Description

90%Weight 40%Scale 1-3

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A tight, third-person description with an explicit 'Use for' trigger clause and strong natural-language trigger coverage. Its only weakness is specificity: it names the diagnosis action and domain rather than enumerating multiple concrete capabilities.

DimensionReasoningScore

Specificity

Names the domain ('bugs and failing behavior') and a concrete action ('Diagnosis loop') but does not list multiple distinct concrete actions the way a score-3 anchor like 'extract text, fill forms, merge documents' does, so it is not comprehensive.

2 / 3

Completeness

Explicitly answers both what ('Diagnosis loop for bugs and failing behavior') and when via an explicit 'Use for ...' trigger clause, satisfying the score-3 anchor.

3 / 3

Trigger Term Quality

Covers natural terms users actually say — 'errors, stack traces, regressions, failed tests, ... debug/fix a bug' — with good breadth including common variations.

3 / 3

Distinctiveness Conflict Risk

Occupies a clear debugging niche with distinct triggers (stack traces, regressions, failed tests, stuck investigations) that are unlikely to fire for unrelated skills.

3 / 3

Total

11

/

12

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
EveryInc/compound-engineering-plugin
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.