CtrlK
BlogDocsLog inGet started
Tessl Logo

diagnose

Diagnose hard bugs and performance regressions with a reproduce-to-regression-test feedback loop. Use when the user reports broken, throwing, failing, or slow behavior or explicitly asks to diagnose or debug it.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

96%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An excellent instruction-only skill: lean, opinionated, assumes competence, and highly actionable, with a rigorously gated six-phase workflow and checklists. The only structural note is that progressive disclosure relies on a single script file while all tactical detail lives in the main body.

DimensionReasoningScore

Conciseness

The body assumes Claude's competence throughout — no explanation of what git bisect, Playwright, or a HAR file is — and every section adds non-obvious judgment ("A 30-second flaky loop is barely better than no loop", "Untagged logs survive; tagged logs die"). Stylistic punches like "Be aggressive. Be creative. Refuse to give up." are heuristics for a perseverance-critical phase, not padding, so it fits the 5 anchor rather than the 4 anchor's 'minor instances of over-explanation'.

5 / 5

Actionability

Ten ranked, concrete loop-construction tactics (curl script, CLI snapshot diff, Playwright, trace replay, `git bisect run` harness, differential loop, `scripts/hitl-loop.template.sh`), a verbatim falsifiable-hypothesis format, a concrete log-tagging convention (`[DEBUG-a4f2]`), and a perf branch with named tools. Per the instruction-only scoring note, absence of code isn't penalized when guidance is this executable; this is copy-paste-ready direction covering the common cases, matching the 5 anchor.

5 / 5

Workflow Clarity

Six clearly sequenced phases with explicit gates ("Do not proceed to Phase 2 until you have a loop you believe in", "Do not proceed until you reproduce the bug"), checklists in Phases 2 and 6, and a closed feedback loop (failing test → fix → watch pass → re-run original scenario). This exceeds the 4 anchor's 'most checkpoints' — validation is explicit at every phase boundary, matching the 5 anchor including error-recovery guidance (no-seam case, non-deterministic-bug escalation path).

5 / 5

Progressive Disclosure

Good structure: clear phase headers, and the one bundle reference (`scripts/hitl-loop.template.sh`) is real, well-signaled with its path, and one level deep. It falls short of the 5 anchor because the entire methodology — including the 10-item loop-construction catalog — is inlined in SKILL.md with no split of deep-detail material into references/, which the anchor expects for a skill of this size. It is clearly above 3, since what's inline is cohesive process guidance rather than reference material dumped into the main file.

4 / 5

Total

19

/

20

Passed

Description

91%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person, concise, with an explicit 'Use when...' clause and a rich set of natural trigger terms. The only weakness is modest overlap risk with general debugging/perf skills, which is inherent to the domain.

DimensionReasoningScore

Specificity

"Diagnose hard bugs and performance regressions with a reproduce-to-regression-test feedback loop" names the domain (hard bugs, perf regressions) and the concrete mechanism (reproduce-to-regression-test loop), with minor gaps — it does not enumerate multiple distinct capabilities the way a 5-anchor example does. It sits clearly above 3, which would cover only 1-2 generic actions with no named methodology.

4 / 5

Completeness

It explicitly answers what ("Diagnose hard bugs and performance regressions with a reproduce-to-regression-test feedback loop") and when ("Use when the user reports broken, throwing, failing, or slow behavior or explicitly asks to diagnose or debug it") with concrete trigger phrases. This matches the 5 anchor verbatim in structure; 4 would require the 'when' clause to be less explicit.

5 / 5

Trigger Term Quality

"broken, throwing, failing, or slow behavior" plus "diagnose or debug" covers the natural synonym family users actually say (it's broken / it's failing / it's slow / help me debug) including the explicit-ask variants. Only peripheral terms like "crash" or "error" are absent, which is not enough to drop to the 4 anchor's "a few natural terms missing".

5 / 5

Distinctiveness Conflict Risk

The framing (reproduce-to-regression-test feedback loop, hard bugs + perf regressions) carves a fairly distinct niche, but "broken/failing/slow behavior" triggers overlap with general coding-fix and performance-profiling skills. Mostly distinct with minor overlap risk against closely related skills — the 4 anchor — rather than the 5 anchor's minimal conflict risk.

4 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

Total

15

/

16

Passed

Repository
EpicenterHQ/epicenter
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.