CtrlK
BlogDocsLog inGet started
Tessl Logo

ce-debug

Diagnosis loop for bugs and failing behavior. Use when asked to debug or fix failing or slow behavior.

66

Quality

80%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./skills/ce-debug/SKILL.md

The canonical home for this skill is ce-debug in EveryInc/compound-engineering-plugin

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An unusually disciplined instruction skill: every phase has explicit gates, the guidance is copy-ready concrete, and detail is genuinely delegated to eight real reference files with imperative read-now signals. The two costs are embedded justificatory prose that inflates token load and Phase 4's routing detail, which sits inline where the other phases push detail to references.

Suggestions

Trim rationale sentences that explain why a rule exists (e.g. "Every option loses something the agent cannot choose on the user's behalf, which is why this question survives"; the "Two facts make it less obvious than it looks" framing) — the rules are enforceable without the justification, saving meaningful tokens in the always-loaded body.

Move the question-2 ships/stays-local decision tree and its PR-capability facts into references/post-fix-handoff.md (or a new routing reference), keeping only the two questions and the three outcomes in the body — this matches how Phases 0-2 and 3 delegate detail.

Collapse the Artifact Root HTML comment markers and validation bullets into the two or three operative lines, or relocate the validation detail to a reference, since it is rarely needed per run.

DimensionReasoningScore

Conciseness

The body assumes Claude's competence throughout (no git/git-hub primers, dense imperative prose) but embeds substantial rationale that could be trimmed, e.g. "Every option loses something the agent cannot choose on the user's behalf, which is why this question survives" and the extended already-pushed-vs-offered explanation in Phase 4's question 2. Not 4 because these over-explanations are more than minor; not 2 because there is no concept-teaching padding and every section drives behavior.

3 / 5

Actionability

Fully concrete instruction-only guidance: exact commands with pitfalls handled ("`git rev-parse --abbrev-ref origin/HEAD` with its `origin/` prefix stripped"), a copy-ready Debug Summary template, enumerated fix-choice options, exact status spellings ("fixed-and-pushed | fixed-not-pushed | diagnosed-no-fix | flaky-infra | needs-human"), and precise delegation rules to named skills. Per the rubric's instruction-only note, the absence of code is not penalized when guidance is this actionable.

5 / 5

Workflow Clarity

Five phases are sequenced in order with explicit validation gates: the causal-chain gate before Phase 3 ("'Somehow X leads to Y' is a gap"), the same-turn presentation-before-gate rule, escalation triggers ("2-3 hypotheses exhausted... or 3 failed fix attempts"), and pre-commit confirmation of unstaged user work. Feedback loops and recovery rules are stated or delegated to named references; nothing in the risky commit/push path lacks a checkpoint.

5 / 5

Progressive Disclosure

All referenced files exist in references/ and are one level deep, clearly signaled with directives ("Read references/investigate.md now and follow it for Phases 0-2"), and the body deliberately keeps only the gates inline. Not 5 because the body is more than an overview: Phase 4's ~20 lines of routing decision logic (ships/stays-local/not-a-git-repo, PR-capability facts) reads as reference-grade detail inlined in SKILL.md rather than split out.

4 / 5

Total

17

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A concise, well-structured description with an explicit 'Use when' clause and natural trigger terms. Its main weakness is thin capability enumeration — the skill actually covers root-cause investigation, test-first fixing, and structured handoff, none of which the description conveys.

Suggestions

Enumerate one or two more concrete actions in the 'what' clause, e.g. "Diagnoses root causes of bugs and failing or slow behavior, then applies test-first fixes" — this lifts specificity toward the several-actions anchors without adding padding.

Add common user synonyms to the trigger clause, such as "error", "broken", "crash", or "regression", to broaden natural-term coverage.

Consider whether "slow behavior" invites overlap with performance-analysis skills; either keep it (if intended) or narrow it to failing behavior to sharpen distinctiveness.

DimensionReasoningScore

Specificity

"Diagnosis loop for bugs and failing behavior" plus "fix failing or slow behavior" names the domain and two concrete actions (diagnose, fix), but doesn't enumerate the fuller capability set (root-cause analysis, regression-test selection, branch/commit handoff). Not 4 because anchor 4 requires several specific actions with only minor coverage gaps; not 2 because it goes beyond a bare domain label.

3 / 5

Completeness

It explicitly answers both parts: what ("Diagnosis loop for bugs and failing behavior") and when ("Use when asked to debug or fix failing or slow behavior") with concrete trigger phrases. Matches the anchor-5 pattern of the good overall examples; a 4 would require the 'when' to be less explicit than it is.

5 / 5

Trigger Term Quality

"debug", "fix", "bugs", "failing", "slow behavior" are all phrases users naturally say. Not 5 because common synonyms like "error", "broken", "crash", or "regression" are absent; not 3 because coverage clearly goes beyond a couple of generic keywords.

4 / 5

Distinctiveness Conflict Risk

Debugging/failing-behavior diagnosis is a clear niche with distinct triggers. Not 5 because "fix failing or slow behavior" is broad enough to overlap with performance-tuning or code-review skills; not 3 because the trigger set is well differentiated from generic "fix code" skills.

4 / 5

Total

16

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
crdant/compound-engineering-plugin
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.