CtrlK
BlogDocsLog inGet started
Tessl Logo

debug-false-positive

Diagnose and fix a Fallow false positive or false negative through extraction, resolution, graph, analysis, reporting, and real-consumer verification.

68

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exemplary lean debugging workflow: tightly sequenced, with genuine validation checkpoints, a self-correcting escalation rule, and zero padding. The only meaningful gap is that the concrete reproduction command and the 'review' invocation are referenced but never shown, which keeps actionability at 4 rather than 5.

Suggestions

Include one example reproduction command (e.g. the actual Fallow CLI invocation with '--format json --quiet') so step 1 is copy-paste ready.

Briefly spell out what 'Run review with the affected surface reviewers' means — which surfaces map to which reviewers — or point to where that mapping is defined.

DimensionReasoningScore

Conciseness

The body is a lean 28-line numbered workflow where every step earns its place: it assumes Claude's competence, never explains what Fallow or static analysis is, and adds only non-obvious operational rules (cache clearing, hypothesis discipline, the 3-failed-fixes stop rule). Matches the 5 anchor 'every token earns its place'.

5 / 5

Actionability

Steps give concrete, executable direction with specifics like '--format json --quiet', clearing the '.fallow/' cache, 'docs/development/quality-gates.md', and 'add a regression test that fails without the fix'. Not 5 because the actual reproduction command is never spelled out (no example invocation of the Fallow binary) and 'Run review with the affected surface reviewers' leaves the reviewer set implicit — minor gaps per the 4 anchor.

4 / 5

Workflow Clarity

The 8-step sequence has explicit validation checkpoints (reproduce before changing code, regression test that must fail without the fix, old-vs-new output comparison on a real consumer) plus a feedback loop for error recovery ('After 3 failed fixes, stop and write down the premise that they share. Question that premise before a fourth fix'). This matches the 5 anchor; not a 4 because no checkpoint is missing or implicit.

5 / 5

Progressive Disclosure

The skill is under 50 lines with no bundle files and no content that belongs in separate files; the single reference ('docs/development/quality-gates.md') points at existing repo documentation, which is an acceptable one-level reference. Per the judging guidelines, a skill under 50 lines with no need for external references scores 5 on well-organized sections, and the numbered-list structure is well-organized.

5 / 5

Total

19

/

20

Passed

Description

70%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A specific, third-person description anchored to a distinct niche (Fallow static-analysis false results), with good natural trigger terms. Its main weakness is the complete absence of a 'when to use' clause, which caps completeness at 3.

Suggestions

Add an explicit trigger clause, e.g. 'Use when Fallow reports a finding that looks wrong or misses a real issue, or when the user mentions a Fallow false positive or false negative.'

Include natural synonyms users might say, such as 'wrong result', 'incorrect finding', or 'missed issue', to broaden trigger-term coverage.

Consider naming the concrete defect surfaces (e.g. extraction, reachability, suppression) in terms users would recognize from Fallow output rather than internal pipeline layer names.

DimensionReasoningScore

Specificity

The description names the tool ('Fallow'), the defect classes ('false positive or false negative'), and enumerates concrete pipeline stages ('extraction, resolution, graph, analysis, reporting, and real-consumer verification'). This lists several specific actions with only minor gaps in coverage, matching the 4 anchor rather than 5 because the stage list describes internal pipeline layers rather than comprehensive user-facing actions.

4 / 5

Completeness

It clearly answers 'what' (diagnose and fix false results through the named stages) but contains no 'Use when...' clause or equivalent explicit trigger guidance, which caps completeness at 3 per the judging guidelines. It is not a 2 because the 'what' is specific and comprehensive, not vague.

3 / 5

Trigger Term Quality

'false positive', 'false negative', 'diagnose', 'fix', and 'Fallow' are phrases a user would naturally say when needing this skill. Not the 5 anchor because common synonyms like 'wrong result', 'incorrect finding', or 'bad report' are missing.

4 / 5

Distinctiveness Conflict Risk

Scoping to 'a Fallow false positive or false negative' creates a clear niche with distinct triggers tied to a named tool and specific defect classes; minimal overlap risk with any generic debugging skill. It cannot score 4 because it is at least as distinct as the PDF-scoped 5 anchor example.

5 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
fallow-rs/fallow
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.