CtrlK
BlogDocsLog inGet started
Tessl Logo

creating-debug-tests-and-iterating

Use this skill when faced with a difficult debugging task where you need to replicate some bug or behavior in order to see what is going wrong.

57

Quality

72%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

High

Do not use without reviewing

Fix and improve this skill with Tessl

tessl review fix ./.agency/plugins/nori/skills/creating-debug-tests-and-iterating/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

61%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers a clear iterative debug-test loop and sensible guidance on real-vs-mock testing, and it is reasonably token-efficient. Its weaknesses are placeholder-only code examples, a missing debug-script example, broken step numbering (3 jumps to 5), and an unreconciled contradiction about when to run other tests.

Suggestions

Fix the step numbering (the list jumps from 3 to 5) and reconcile the contradiction between 'ignore other tests until you have a working example' and step 7 'Make sure other tests pass'.

Include one minimal executable debug-script example instead of only placeholder paths like './path/to/cli.sh', and show a concrete API call sample for the server case.

Remove the fake <system-reminder> injections and the duplicated 'ignore other tests' instruction; fold them into a single explicit step to tighten conciseness.

DimensionReasoningScore

Conciseness

The body is lean with no padding or explanations of concepts Claude already knows, and code examples are compact. Minor tightening is possible: the closing note "Do NOT get in a loop where you just keep running other tests" repeats the earlier "ignore any existing tests until you have a working example", fitting 'Efficient; minor instances of over-explanation that could be trimmed' rather than the fully lean anchor 5.

4 / 5

Actionability

Concrete elements exist (bash/Python/Node snippets for CLI calls, loop steps), but they use placeholder paths like "./path/to/cli.sh" and the API section only says "Call to the server using scripting language of choice" with no example. There is no example of an actual debug script — the core deliverable — matching 'Some concrete guidance but incomplete; pseudocode instead of executable code; missing key details'.

3 / 5

Workflow Clarity

A real feedback loop is specified (add logs → run script → analyze → update, repeated until fixed), but the numbered list skips from step 3 to step 5, and instructions contradict each other: "ignore any existing tests until you have a working example" vs. step 7 "Make sure other tests pass", with no reconciliation of when the switch happens. This fits 'Steps listed but validation gaps; sequence present but checkpoints missing or implicit' rather than the mostly-checkpointed anchor 4.

3 / 5

Progressive Disclosure

The skill is short (<50 lines) with no bundle files, and sections are reasonably organized with one clearly signaled external reference ("read the .claude/skills/webapp-testing/SKILL.md"). Under the simple-skill guidance this could score 5, but the unnumbered-to-body structure (the <required> block restating steps that later sections cover) leaves minor organization gaps, placing it at anchor 4.

4 / 5

Total

14

/

20

Passed

Description

70%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concise, uses an explicit 'Use when' clause, and names a clear niche (replicating bugs to diagnose what's wrong). Its main gaps are limited action coverage and a single, somewhat generic trigger condition with missing common synonyms.

Suggestions

List the concrete actions the skill covers (e.g. 'Write external debug scripts, add logs, iterate until the bug is fixed') to raise specificity.

Broaden trigger terms with natural synonyms such as 'reproduce', 'repro', or 'debug a failing behavior' users would say.

Make the 'when' more concrete, e.g. 'Use when the user asks to reproduce, replicate, or diagnose a bug or unexpected behavior.'

DimensionReasoningScore

Specificity

"replicate some bug or behavior in order to see what is going wrong" names the debugging domain and one concrete action (replicating the bug), but coverage is not comprehensive — no mention of writing test scripts, iterating, or cleanup. This matches the anchor 'Names domain and 1-2 concrete actions, but not comprehensive'; score 4 would require several listed actions.

3 / 5

Completeness

Both parts are present: the 'what' ("replicate some bug or behavior in order to see what is going wrong") and an explicit 'when' ("Use this skill when faced with a difficult debugging task"). The 'when' is a single general condition rather than concrete trigger phrases (e.g. 'when the user asks to reproduce a bug'), so it sits at anchor 4 rather than 5.

4 / 5

Trigger Term Quality

Natural terms like "difficult debugging task", "replicate", and "bug" are phrases users would actually say. A few common synonyms are missing ("reproduce", "repro", "troubleshoot"), which fits 'Good keyword coverage; a few natural terms missing' rather than the comprehensive anchor 5.

4 / 5

Distinctiveness Conflict Risk

Bug replication for debugging is a fairly distinct niche with clear triggers, though it could overlap with general testing or test-writing skills. 'Mostly distinct; minor overlap risk with closely related skills' is the best fit — anchor 5 would require a clearer niche with fully distinct triggers.

4 / 5

Total

15

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
microsoft/FluidFramework
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.