CtrlK
BlogDocsLog inGet started
Tessl Logo

fix-bug

Resolves a single bug from any starting evidence — Dash0 telemetry (span, log, web event, RUM error link), a stack trace, an error message, a code pointer (file:line), a screen recording, a Linear ticket URL, or a free-text symptom. Classifies the input, triages complexity (Phase 0.5) to pick a fast lane or a full holistic-analysis lane, runs a pre-flight sweep, locks a failing reproduction (delegating to /tdd, /e2e-testing, or /e2e-testing-mobile by layer), delegates root-cause analysis to the isolated rca-investigator agent on complex bugs, gates on confidence(analysis), and at >= 92 % hands off without human confirmation — fast lane via aw-create-plan + aw-executor, standard lane via aw-planner + aw-executor, both under a CEGIS refinement contract. A bug-fix-verifier agent grades the PR before undrafting; for telemetry-sourced bugs an optional Phase 8 polls the originating signal post-deploy. --analyse-only stops at the proposal; --force-holistic skips the fast lane. Triggers on "/fix-bug".

65

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

70%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-engineered orchestrator spec: phase sequencing, mechanical gates, and error-recovery loops are exemplary, and direct instructions are copy-paste concrete. It loses points on token efficiency (the Architecture/Modes/Phase 5–6/Risks/Key Principles sections repeat the same rules four to five times) and on progressive disclosure as delivered — the skill is designed as a thin index over rules/*.md and templates/*.md files that are not present in this bundle, breaking the delegation chain that its most important procedures depend on.

Suggestions

Ship the rules/*.md and templates/*.md files in the bundle (or inline their critical content): the body delegates ten core procedures — complexity triage, evidence resolution, preflight, reproduction, handoff, verification — to rule files that are absent, leaving the pipeline unexecutable as delivered.

Collapse the duplicated threshold/lane statements: the 92% auto-implement rule, lane triggers, and no-force-proceed rules each appear in the Architecture diagram, Modes table, Phase 5, Phase 6, Risks, and Key Principles — keep one authoritative statement plus the phase diagram and delete the rest.

Trim the oversized structured-error blocks (e.g. the Step 5a reproduction-gate error) to the essential failure message plus a pointer to the governing rule file; the full enumerated recovery text is repeated context the rule file already carries.

DimensionReasoningScore

Conciseness

The operational core (tables, regexes, commands, gates) is dense and information-rich, but the ~810-line body restates the same lane/threshold/gate rules repeatedly — the 92% auto-implement rule appears in the Architecture diagram, the Modes table, Phase 5, the Risks table, and Key Principle 6, and the 12-item Key Principles section largely re-summarises earlier sections ("Auto-implement at ≥ 92 % is fully autonomous. No human confirmation."). This fits anchor 3 — mostly efficient but with clear tightening opportunities — and not 2, since almost none of it explains concepts Claude already knows.

3 / 5

Actionability

Highly executable where it speaks directly: exact detection regexes ("Matches `https?://linear\.app/.+/issue/`"), exact commands ("gh pr view --json url,isDraft,headRefName", "git rev-parse --abbrev-ref HEAD"), verbatim fail-closed error blocks, closed best-effort reason lists, and a copy-paste output template. It is not 5: several key procedures defer to files absent from this bundle — "Walk the 14-row signal table in rules/complexity-triage.md", evidence-resolution, and reproduction layer routing all live in rules/*.md that are not shipped — so an executor following only what is written hits real gaps.

4 / 5

Workflow Clarity

Phases 0–8 are explicitly sequenced with three mechanical validation gates (Step 5a reproduction gate, Step 6.pre protected-branch assertion, Step 7.pre draft-PR assertion), each with a deterministic procedure and fail-closed structured error, plus feedback loops (route back to Phase 2.5, CEGIS 3-round cap with standard-lane fallback, simple→complex triage upgrade). This matches the anchor-5 pattern of validate → fix → retry with checklists for a complex process.

5 / 5

Progressive Disclosure

Signaling is excellent — "This SKILL.md is a thin index... Load only what the current phase asks for", per-phase Rule/Template loading tables, one-level-deep references — but scored against the actual bundle: of the 13 referenced paths only references/research-sources.md exists; the 10 rules/*.md and 2 templates/*.md files the body depends on are absent, so the index points at missing files. Not 4 or 5: a thin index whose target files are not present cannot count as well-organized disclosure; not 2: the in-body structure is clear and every reference is systematically tabulated rather than buried.

3 / 5

Total

15

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An unusually strong description: it states what the skill does at every pipeline stage in concrete, third-person action verbs, explicitly declares its trigger (the /fix-bug command) and the full set of accepted input shapes, and occupies a distinct niche. The only weakness is mild — a few natural trigger synonyms (crash, exception, regression) are missing, and the density of internal mechanism names (aw-planner, CEGIS contract) adds length without user-facing value.

DimensionReasoningScore

Specificity

The description enumerates a full pipeline of concrete actions — "Classifies the input, triages complexity (Phase 0.5)", "locks a failing reproduction (delegating to /tdd, /e2e-testing, or /e2e-testing-mobile by layer)", "gates on confidence(analysis)", "bug-fix-verifier agent grades the PR before undrafting", "polls the originating signal post-deploy" — comprehensive coverage of every phase. It is not 4: there are no meaningful coverage gaps; each capability is named as a specific action rather than a generic verb.

5 / 5

Completeness

Both questions are answered explicitly: what — "Resolves a single bug from any starting evidence" followed by the complete action chain through verification and telemetry polling; when — "Triggers on \"/fix-bug\"" plus the enumerated trigger inputs (Dash0 telemetry, stack trace, error message, code pointer, screen recording, Linear ticket, free-text symptom), which is equivalent explicit trigger guidance, so the missing-'Use when' cap at 3 does not apply. It is not 4: the 'when' is concrete and explicit, not merely implied or underspecified.

5 / 5

Trigger Term Quality

Natural terms a bug-reporting user would actually say are present — "stack trace", "error message", "screen recording", "Linear ticket", "code pointer (file:line)", "free-text symptom", "bug" — plus the explicit "/fix-bug" command trigger. It is not 5: common synonyms such as "crash", "exception"/"traceback", "regression", or "failing test" are absent; it is not 3: the included keywords directly match user phrasing rather than only technical jargon.

4 / 5

Distinctiveness Conflict Risk

A clear niche — single-bug resolution from a specific evidence set with distinct triggers ("Dash0 telemetry", "Linear ticket URL", "/fix-bug", "rca-investigator") that would not collide with general coding or document skills. It is not 4: no meaningful overlap risk with closely related skills is identifiable from the description.

5 / 5

Total

19

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (876 lines); consider splitting into references/ and linking

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 45 missing, 6 suspicious

Warning

Total

12

/

16

Passed

Repository
mthines/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.