CtrlK
BlogDocsLog inGet started
Tessl Logo

benchmark-fp-fn-audit

Audit React Doctor against ReactBench or similar diagnostic benchmark corpora for confirmed false positives, false negatives, taxonomy gaps, and verifier artifacts. Use when analyzing rd.log, rd-before.json, rd-after.json, model.patch, result.json, reward/test logs, rule distributions, or when asked to perform a second adversarial pass over React Doctor benchmark findings.

68

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A detailed, well-sequenced audit workflow with concrete artifact paths, classification rules, and output schemas. Main weaknesses are the lack of an explicit output-validation feedback loop and a missing pointer to the existing reference file.

Suggestions

Link to references/benchmark-artifacts.md from the body (e.g., 'See [benchmark-artifacts.md](references/benchmark-artifacts.md) for the evidence map') instead of inlining the trial-lead list and corpus path, which duplicates the reference.

Add an explicit output-validation checkpoint in step 5, e.g. re-load the JSONL and confirm every record's cited trial/file/lines exist and classifications match the evidence rules before finalizing.

Tighten restatements of expert-known facts (e.g., the 'react_doctor=1 is a gate result' aside) to lift conciseness toward the lean anchor.

DimensionReasoningScore

Conciseness

Largely efficient with concrete rule names, artifact paths, and a JSONL schema that earn their tokens; a few sentences restate expert knowledge (e.g., 'react_doctor=1 is a gate result, not proof'), keeping it just below the lean anchor 5.

4 / 5

Actionability

Provides concrete commands ('rg --files'), exact artifact paths, a full JSONL field schema with enum values, and a prioritization formula, but as an instruction-only skill leaves some execution mechanics to Claude's expertise.

4 / 5

Workflow Clarity

A clearly sequenced five-step workflow with verification-style checkpoints ('recompute before relying on a prior summary', 'find a nearby true-positive counterexample'), though it lacks an explicit validate-fix-retry loop on its outputs.

4 / 5

Progressive Disclosure

Well-structured with clear section headers, but the available references/benchmark-artifacts.md is never linked or signaled from the body, and trial-lead content that overlaps the reference is inlined rather than delegated.

3 / 5

Total

15

/

20

Passed

Description

95%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, specific description with concrete trigger filenames and an explicit 'Use when' clause that clearly delimits when the skill applies. The only minor gap is that the action side is framed as a single audit verb decomposed into finding types rather than a list of distinct actions.

DimensionReasoningScore

Specificity

Names the audit action and decomposes it into concrete finding types ('confirmed false positives, false negatives, taxonomy gaps, and verifier artifacts'), but does not list multiple distinct action verbs, so it sits just below the comprehensive anchor 5.

4 / 5

Completeness

Explicitly answers both 'what' (audit for FP/FN/taxonomy gaps/verifier artifacts) and 'when' via a concrete 'Use when analyzing... or when asked to perform a second adversarial pass' clause.

5 / 5

Trigger Term Quality

Comprehensive natural trigger terms including exact artifact filenames and extensions ('rd.log', 'rd-before.json', 'rd-after.json', 'model.patch', 'result.json', 'reward/test logs', 'rule distributions') that a user would naturally mention.

5 / 5

Distinctiveness Conflict Risk

Highly niche ('React Doctor', 'ReactBench', specific trial artifacts) with distinct triggers and minimal overlap risk with other skills.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
millionco/react-doctor
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.