CtrlK
BlogDocsLog inGet started
Tessl Logo

tooluniverse-diagnostic-test-evaluation

Diagnostic test / biomarker accuracy — sensitivity, specificity, PPV, NPV, likelihood ratios, accuracy from a 2x2 table; ROC curve, AUC, and the optimal cutoff (Youden) for a continuous biomarker; and post-test probability via Bayes. Use when you have test results vs a gold standard (binary 2x2, or a continuous score + true labels) and need to judge how good the test is, pick a threshold, or compute the probability of disease given a result. Emphasizes the prevalence-dependence of PPV/NPV.

72

Quality

90%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, highly actionable body: every step has an executable command, a triage table routes by input shape, and the bundle script is real and consistent with its documented usage. Weak spots are mild: some textbook formulas and AUC bands re-taught inline, repeated prevalence warnings, and no explicit validation checkpoints in the operational flow.

Suggestions

Trim the 2x2 formula table and AUC interpretation bands (or move them to a short reference file) — Claude already knows these; keep only the prevalence-dependence column and the PPV/NPV trap, which are the skill's real value-add and currently stated three times.

Add one validation checkpoint per step: verify the 2x2 cell counts are consistent/non-negative before quoting metrics, and note that roc_analysis.py exits with an error if the CSV lacks both label classes (and silently skips malformed rows).

Consolidate the prevalence-dependence guidance (metric-table column, the 'PPV/NPV trap' callout, and the first gotcha) into a single statement to remove repetition and recover token budget.

DimensionReasoningScore

Conciseness

Mostly efficient, structured, and non-padded, but the 2x2 metric formula table (TP/(TP+FN), TN/(TN+FP), etc.) and the AUC interpretation bands restate knowledge Claude already has, and the prevalence-dependence point is made three times (metric table, 'PPV/NPV trap' callout, and gotchas).

4 / 5

Actionability

Fully executable, copy-paste-ready commands for every path: 'tu run Epidemiology_diagnostic '{"operation":"diagnostic","tp":90,...}'' with concrete numbers, 'python skills/.../scripts/roc_analysis.py --input scores.csv' with the CSV column spec, and the ROC_analysis tool form with both inline-array and CSV variants.

5 / 5

Workflow Clarity

The triage table ('A 2x2 table at a fixed cutoff -> Step 1', 'A continuous biomarker score + true labels -> Step 2') plus explicit cross-step chaining ('Once you choose a cutoff, build its 2x2 and run Step 1') give a clear sequence, but there are no explicit validation/error-recovery checkpoints (e.g., sanity-checking the 2x2 counts or handling a malformed CSV).

4 / 5

Progressive Disclosure

Good structure with well-organized sections and a single script bundle whose body reference matches the actual file at scripts/roc_analysis.py (usage, CSV columns, and outputs are consistent). Minor gaps: the ~90-line body keeps formula tables and gotcha reference material inline that could live in a separate reference file, and no references/ files exist.

4 / 5

Total

17

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete, comprehensive capability listing in third person, an explicit 'Use when...' trigger clause covering both binary 2x2 and continuous-biomarker inputs, and a well-differentiated niche. Only minor keyword coverage gaps (missing 'confusion matrix' and spelled-out PPV synonyms) keep trigger terms at 4.

DimensionReasoningScore

Specificity

Lists multiple specific concrete actions across all three sub-capabilities — 'sensitivity, specificity, PPV, NPV, likelihood ratios, accuracy from a 2x2 table; ROC curve, AUC, and the optimal cutoff (Youden) for a continuous biomarker; and post-test probability via Bayes' — with comprehensive coverage and no gaps.

5 / 5

Completeness

Clearly answers both what (three concrete capability sets) and when, with an explicit 'Use when you have test results vs a gold standard (binary 2x2, or a continuous score + true labels) and need to judge how good the test is, pick a threshold, or compute the probability of disease given a result' trigger clause.

5 / 5

Trigger Term Quality

Good coverage of natural terms users would say — 'diagnostic test', 'sensitivity', 'specificity', 'PPV', 'ROC', 'AUC', 'biomarker', 'gold standard', 'pick a threshold' — but misses common synonyms such as 'confusion matrix', spelled-out 'positive predictive value', or 'false positive/negative rate'.

4 / 5

Distinctiveness Conflict Risk

A clear niche — diagnostic test/biomarker accuracy at a fixed cutoff, across cutoffs (ROC), and post-test probability — with distinct triggers; only minor overlap risk with a general epidemiology skill.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
mims-harvard/ToolUniverse
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.