CtrlK
BlogDocsLog inGet started
Tessl Logo

debug-false-positive

Diagnose and fix a Fallow false positive or false negative through extraction, resolution, graph, analysis, reporting, and real-consumer verification.

68

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exemplary lean debugging workflow: a numbered, well-sequenced procedure with genuine validation checkpoints (reproduce-first, failing regression test, real-consumer verification) and an anti-rut feedback loop after three failed fixes. The only minor gap is that a few steps rely on directive phrasing rather than exact commands.

DimensionReasoningScore

Conciseness

The body is ~27 lines of dense, imperative instructions with zero padding and no explanation of concepts Claude already knows ("Change no code before this command exists", "Fix the pattern, not the instance"). Every token earns its place, matching the 'lean and efficient; assumes Claude's competence' anchor.

5 / 5

Actionability

Concrete details are present: the reproduce command flags ("--format json --quiet"), the cache path (".fallow/"), a pointer to test evidence ("docs/development/quality-gates.md"), and quantified directives ("3 to 5 ranked hypotheses", "After 3 failed fixes"). A few steps (reduce to a minimal fixture, trace through the pipeline layers) are directives without an exact command, which is the minor gap that separates this from anchor 5.

4 / 5

Workflow Clarity

An explicit 8-step sequence with validation checkpoints throughout: reproduce before changing code (step 1), a regression test that must fail without the fix (step 6), real-consumer output comparison (step 7), and a review pass (step 8). It even includes an error-recovery feedback loop — "After 3 failed fixes, stop and write down the premise that they share. Question that premise before a fourth fix" — matching the top anchor's requirements.

5 / 5

Progressive Disclosure

The skill is under 50 lines, has no bundle files, and needs no external references beyond a single well-signaled repo path (docs/development/quality-gates.md) used to justify a practice rather than park content. Per the rubric's simple-skill guidance, a short body with well-organized numbered structure scores 5; content is appropriately inline with nothing that belongs in a separate file.

5 / 5

Total

19

/

20

Passed

Description

70%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, specific description that clearly states what the skill does across the Fallow pipeline, with natural trigger terms and a distinct niche. Its main weakness is the absence of any "Use when..." clause, which caps completeness and leaves trigger conditions implicit.

Suggestions

Add an explicit 'Use when' clause, e.g., "Use when Fallow reports a wrong or unexpected finding, or when a user says a result is a false positive or false negative."

Include common user-side synonyms and variations such as 'wrong result', 'incorrect finding', or 'misreport' alongside 'false positive / false negative' to broaden natural trigger coverage.

Slightly sharpen the action list by stating outcomes (e.g., 'locate the responsible pipeline layer and add a regression test') so the capabilities read as concrete actions rather than pipeline-stage names.

DimensionReasoningScore

Specificity

The description names the domain ("Fallow false positive or false negative") and enumerates six concrete workflow actions ("extraction, resolution, graph, analysis, reporting, and real-consumer verification"), giving broad but not fully comprehensive coverage. It sits above anchor 3 (only 1-2 concrete actions) but the actions are pipeline-stage names rather than fully spelled-out operations, so it does not clearly match the comprehensiveness of anchor 5.

4 / 5

Completeness

The "what" is clear — diagnose and fix false positives/negatives through the listed pipeline stages — but there is no "when" clause or equivalent explicit trigger guidance (e.g., "Use when Fallow reports..."). Per the judging guidelines, a missing 'Use when...' clause caps completeness at 3.

3 / 5

Trigger Term Quality

"Diagnose and fix", "false positive", and "false negative" are natural phrases a user troubleshooting Fallow output would actually say. Coverage is good but missing common variations and synonyms (e.g., "wrong result", "incorrect finding", "misreport"), which keeps it below anchor 5.

4 / 5

Distinctiveness Conflict Risk

"Fallow false positive or false negative" carves out a clear niche tied to a specific tool's accuracy debugging, with distinct triggers unlikely to collide with other skills. It is not below 5 because the conflict risk is genuinely minimal, and it is well above 4's 'minor overlap risk'.

5 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
fallow-rs/fallow
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.