CtrlK
BlogDocsLog inGet started
Tessl Logo

tooluniverse-diagnostic-test-evaluation

Diagnostic test / biomarker accuracy — sensitivity, specificity, PPV, NPV, likelihood ratios, accuracy from a 2x2 table; ROC curve, AUC, and the optimal cutoff (Youden) for a continuous biomarker; and post-test probability via Bayes. Use when you have test results vs a gold standard (binary 2x2, or a continuous score + true labels) and need to judge how good the test is, pick a threshold, or compute the probability of disease given a result. Emphasizes the prevalence-dependence of PPV/NPV.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

The canonical home for this skill is tooluniverse-diagnostic-test-evaluation in mims-harvard/ToolUniverse

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, highly actionable skill body: every step is executable with real commands, the case-routing table makes the workflow unambiguous, and the domain gotchas (prevalence dependence, spectrum bias, single-cutoff optimism) are genuinely valuable. The main improvement room is trimming standard textbook content (metric formula table, AUC interpretation bands) that Claude already knows.

DimensionReasoningScore

Conciseness

Largely efficient — the 'PPV/NPV trap' callout, prevalence-dependence column, Gotchas, and Honest limitations are non-obvious content that earns its tokens. But the sensitivity/specificity/PPV/NPV/LR formula-and-definition table and the AUC interpretation bands (0.7–0.8 acceptable, etc.) re-state standard epidemiology Claude already knows and could be trimmed, matching 'efficient; minor instances of over-explanation'.

4 / 5

Actionability

Fully executable throughout: a complete 'tu run Epidemiology_diagnostic' command with real parameters, both inline-array and CSV forms of ROC_analysis, the bundled script invocation with documented CSV columns, and a concrete Bayes example with a worked numeric result — copy-paste ready and covering all three common cases.

5 / 5

Workflow Clarity

The 'Which case are you in?' routing table gives an explicit decision path, Steps 1–3 are clearly sequenced with cross-links ('build its 2×2 and run Step 1', 'plug the true prevalence in'), and each step is a single unambiguous action. No destructive or batch operations exist, so no validation checkpoint is required.

5 / 5

Progressive Disclosure

Good structure: the single bundle file (scripts/roc_analysis.py, verified to exist and match the documented interface) is clearly signaled and one level deep, and sections are well organized. At ~90 lines the body carries more inline reference material (two metric tables, dual ROC tool/script paths) than a lean overview would, keeping it just below the top anchor.

4 / 5

Total

18

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An excellent description: comprehensive and specific on capabilities, with an explicit, input-shape-aware 'Use when' trigger clause and rich natural trigger vocabulary. Its only weakness is slight breadth around 'biomarker accuracy' that could overlap with adjacent statistics/epidemiology skills.

DimensionReasoningScore

Specificity

The description enumerates concrete actions across all three modes — 'sensitivity, specificity, PPV, NPV, likelihood ratios, accuracy from a 2x2 table', 'ROC curve, AUC, and the optimal cutoff (Youden)', and 'post-test probability via Bayes' — comprehensive coverage matching the anchor for multiple specific concrete actions, in proper third-person voice.

5 / 5

Completeness

Both 'what' (the three analysis capabilities) and 'when' are explicit and concrete: 'Use when you have test results vs a gold standard (binary 2x2, or a continuous score + true labels) and need to judge how good the test is, pick a threshold, or compute the probability of disease given a result.' The 'when' clause specifies input shapes and user intents, matching the top anchor.

5 / 5

Trigger Term Quality

Natural user vocabulary is thoroughly covered: 'diagnostic test', 'gold standard', 'sensitivity', 'specificity', 'PPV', 'NPV', 'ROC', 'AUC', 'cutoff', 'threshold', 'biomarker', 'probability of disease given a result'. Both binary and continuous-biomarker phrasings users would actually say are present.

5 / 5

Distinctiveness Conflict Risk

The 'test results vs a gold standard' framing carves a clear niche, but 'biomarker accuracy' and 'pick a threshold' have minor overlap risk with closely related statistical-modeling and epidemiological-analysis skills — fits 'mostly distinct; minor overlap risk' rather than the fully distinct top anchor.

4 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
mims-harvard/ToolUniverse
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.