CtrlK
BlogDocsLog inGet started
Tessl Logo

diagnosing-bugs

面向棘手缺陷和性能回退的诊断循环。适用于用户说 “diagnose” / “debug this”,或报告某些东西 broken、throwing、failing、slow 时。

68

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An unusually disciplined process skill: a clearly sequenced six-phase workflow with explicit validation checkpoints, checklists, and feedback loops, plus fully actionable concrete artifacts (hypothesis format, debug-tag convention, a real HITL script). The only notable gaps are a handful of trimmable coaching sentences and the absence of any one-level-deep reference files for the long enumeration sections.

Suggestions

Move the ten 'Ways to construct one' methods (or the perf-branch guidance in Phase 4) into a reference file such as references/loop-methods.md, keeping a short ranked summary in SKILL.md with clearly signaled links — this would push progressive disclosure toward the ideal overview-plus-references shape.

Trim purely motivational sentences (e.g. "要强硬、要有创造力、拒绝放弃", "bug 就修好了 90%") or compress them into the instructional lines they reinforce, saving tokens without losing the discipline.

Consider adding a brief opening line stating the phase map (1 loop → 2 reproduce → 3 hypothesise → 4 instrument → 5 fix → 6 cleanup) so the overall sequence is graspable before diving into Phase 1's long core section.

DimensionReasoningScore

Conciseness

The body is dense and imperative throughout — it never explains concepts Claude already knows (no explanation of what git bisect or Playwright is) and mostly every line instructs: "给每条 debug log 加唯一 prefix,例如 [DEBUG-a4f2]", "一次只改变一个变量". A few coaching/emphasis sentences could be trimmed ("要强硬、要有创造力、拒绝放弃", "这就是这个 skill 的核心", "一个 30 秒且 flaky 的 loop 几乎不比没有 loop 好" is borderline motivational rather than instructional), so it sits at anchor 4 (efficient, minor instances that could be trimmed) rather than anchor 5's every-token-earns-its-place.

4 / 5

Actionability

For an instruction-only skill the guidance is fully actionable: a ranked list of ten concrete loop-construction methods each naming specific tools ("Headless browser script(Playwright / Puppeteer)", "git bisect run", "performance.now()"), a copy-paste hypothesis format template ("If <X> is the cause, then <changing Y> will make the bug disappear"), an executable debug-tag convention ([DEBUG-a4f2] with grep-based cleanup), and a real, complete, runnable script at scripts/hitl-loop.template.sh (verified to exist and be executable). This matches anchor 5's 'specific examples cover the common cases'; it exceeds anchor 4 because the concrete artifacts are complete rather than having gaps — situation-dependent steps (probe placement) are inherently variable, which the rubric accepts.

5 / 5

Workflow Clarity

Six explicitly numbered phases, each with an explicit completion criterion or checklist: Phase 1's red-capable/deterministic/fast/agent-runnable checklist, Phase 2's "每个剩余元素都是 load-bearing", Phase 5's fail-then-pass regression sequence, and Phase 6's cleanup checklist ("所有 [DEBUG-...] instrumentation 已移除(grep prefix)"). Validation and feedback loops (red/green signal, re-run original repro, re-validate) are pervasive, matching anchor 5 (clear sequence, explicit validation, feedback loops, checklists) and clearly above anchor 4's 'most checkpoints'.

5 / 5

Progressive Disclosure

Structure is good: the body is organized into clear phase sections, and its one bundle reference — "用 scripts/hitl-loop.template.sh 驱动人" — is clearly signaled and points to a real file in the bundle. However, at ~130 lines the SKILL.md inlines content that could arguably live in reference files (the ten construction methods, the perf-branch guidance), and there is no top-level overview/navigation to optional deeper material, so it lands at anchor 4 (good structure, references mostly clear, minor organization gaps) rather than anchor 5's clear overview with well-signaled one-level-deep references.

4 / 5

Total

18

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A compact, well-constructed description that clearly states both what the skill does and when to use it, with genuinely natural trigger phrasing. Its main weakness is that the 'what' stays at the level of a single abstract action ("diagnostic loop") without naming the concrete steps the skill actually guides.

Suggestions

Enumerate 2-3 concrete actions in the 'what' clause, e.g. "构建 tight 的 reproduction loop、bisect、instrument、写 regression test 的诊断循环" (a diagnostic loop that builds tight reproduction loops, bisects, instruments, and writes regression tests).

Add one or two common symptom synonyms to the trigger list — e.g. "crash"、"error"、"not working"、"regression" — to catch phrasings users say that the current list misses.

Consider adding a brief qualifier that the skill targets non-obvious/tricky bugs specifically, which would sharpen distinctiveness against a generic code-fixing skill.

DimensionReasoningScore

Specificity

The description names the domain concretely ("棘手缺陷和性能回退" — tricky defects and performance regressions) and one action ("诊断循环" — a diagnostic loop), but does not enumerate the concrete sub-actions the skill actually performs (building reproduction loops, bisecting, instrumenting, fixing with regression tests). This matches anchor 3 ('names domain and 1-2 concrete actions, but not comprehensive'); it is above anchor 2 because the domain and action are specific rather than generic, but below anchor 4 because only a single action is named.

3 / 5

Completeness

It explicitly answers both questions: the 'what' is "面向棘手缺陷和性能回退的诊断循环" (a diagnostic loop for tricky defects and performance regressions), and the 'when' is explicit with concrete trigger phrases — "适用于用户说 'diagnose' / 'debug this',或报告某些东西 broken、throwing、failing、slow 时" (use when the user says diagnose/debug this, or reports something broken, throwing, failing, slow). This matches anchor 5 (both what AND when with concrete trigger phrases) and exceeds anchor 4, whose 'when' clause is more generic.

5 / 5

Trigger Term Quality

Natural trigger phrases a user would actually say are present: "diagnose" / "debug this", plus symptom reports "broken、throwing、failing、slow". Good coverage, but common variations like "crash", "error", "not working", "regression", or "stack trace" are missing, so it stops short of the comprehensive synonym coverage of anchor 5 while clearly exceeding the sparse keyword coverage of anchor 3.

4 / 5

Distinctiveness Conflict Risk

The trigger set (diagnose/debug/broken/throwing/failing/slow) carves out a recognizable debugging niche that is unlikely to fire for e.g. document or data-analysis skills. However, broad symptom words like "broken" and "slow" could overlap with a general code-fixing or performance-tuning skill, matching anchor 4 (mostly distinct, minor overlap risk with closely related skills) rather than anchor 5's minimal conflict risk.

4 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
vinvcn/mattpocock-skills-zh-CN
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.