CtrlK
BlogDocsLog inGet started
Tessl Logo

tooluniverse-variant-predictor-dms-validation

Validate a variant-effect predictor (AlphaMissense, ESM-C SAE, ESM logits, EVE, conservation scores, or any per-variant numeric score) against experimental deep mutational scanning (DMS) data. Computes per-variant predictor scores, splits variants into neutral vs disruptive groups by DMS effect, runs a Mann-Whitney U test on the predictor scores, and sweeps the stratification thresholds for robustness. Use when you need to know whether a predictor's scores track real functional disruption on a specific protein.

68

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

The canonical home for this skill is tooluniverse-variant-predictor-dms-validation in mims-harvard/ToolUniverse

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers an excellent, executable workflow with a standout mandatory validation gate and feedback loops, and nearly all code is runnable as written. Its main weaknesses are a monolithic structure with no progressive disclosure (no reference files despite ~440 lines, several of which are predictor-specific detail that belongs in separate files) and minor variable-naming inconsistencies that break copy-paste readiness in Step 6.

Suggestions

Split predictor options A-D into separate reference files (e.g. references/predictors-alphamissense.md) and keep SKILL.md as the workflow overview with one-line pointers, so the core 6-step procedure stays lean.

Fix the variable mismatch between Step 3.5/4 ('s_neutral_finite', 's_disruptive_finite') and Step 6's plotting code ('s_neutral', 's_disruptive') so the visualization block is copy-paste ready.

Trim inline verbosity: condense the full AlphaMissense response example and the 'eval that motivated this skill' anecdote, and either remove or properly locate the reference to 'tests/integration/test_dms_pipeline_e2e_kras.py', which is not part of the skill bundle.

DimensionReasoningScore

Conciseness

The body is dense with non-obvious domain knowledge (tool signatures, Forge cost tables, sign conventions, API response shapes) and mostly avoids explaining concepts Claude already knows, but sections like the full AlphaMissense response dump and the 'eval that motivated this skill' anecdote could be trimmed — fitting the 4-anchor 'minor instances of over-explanation' rather than the 5-anchor's every-token-earns-its-place.

4 / 5

Actionability

Nearly all guidance is executable, copy-paste-ready code with real tool calls and parameters, but there are minor gaps: Step 6's plotting code references 's_neutral'/'s_disruptive' while Steps 3.5-4 define 's_neutral_finite'/'s_disruptive_finite', and option C leaves the log-odds computation to the reader — matching the 4-anchor 'concrete code with minor gaps' rather than fully copy-paste-ready.

4 / 5

Workflow Clarity

Steps 1-6 are clearly sequenced with an explicitly MANDATORY validation checkpoint (Step 3.5's NaN coverage hard gate that raises ValueError with 'INVESTIGATE THE PREDICTOR COMPUTATION STEP before continuing'), error-recovery feedback ('debug the predictor computation step (re-run, check API keys, check log files)'), a sign-convention double-check, and a robustness sweep — matching the 5-anchor's explicit validation steps and feedback loops.

5 / 5

Progressive Disclosure

The 440-line body is well-sectioned but entirely monolithic: no bundle files exist (no references/, scripts/, or assets/), and content that clearly belongs in separate files — the ~80-line AlphaMissense option B with bin-parsing code, the four predictor options, the interpretation tables — is all inlined; it also cites 'tests/integration/test_dms_pipeline_e2e_kras.py', which is not present in the bundle. This fits the 3-anchor 'content that should be separate is inline' rather than the 4-anchor's appropriately split structure.

3 / 5

Total

16

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states a concrete multi-action pipeline, names the specific predictors it covers, and gives an explicit 'Use when' trigger. The only weaknesses are a few missing natural synonyms ('missense', 'benchmark', 'correlate') and slight scope breadth that leaves minor overlap risk with closely related variant-interpretation skills.

DimensionReasoningScore

Specificity

The description lists multiple concrete, comprehensive actions — 'Computes per-variant predictor scores, splits variants into neutral vs disruptive groups by DMS effect, runs a Mann-Whitney U test on the predictor scores, and sweeps the stratification thresholds for robustness' — going beyond the 4-anchor's 'minor gaps in coverage' by enumerating the full pipeline.

5 / 5

Completeness

Both questions are explicitly answered: the 'what' is the enumerated compute/stratify/test/sweep pipeline, and the 'when' is concrete — 'Use when you need to know whether a predictor's scores track real functional disruption on a specific protein.'

5 / 5

Trigger Term Quality

Good natural-term coverage — 'AlphaMissense', 'ESM-C SAE', 'EVE', 'deep mutational scanning (DMS)', "predictor's scores track real functional disruption" — but a few phrases practitioners naturally say ('missense', 'benchmark', 'correlate', 'MaveDB') are absent, matching the 4-anchor rather than the 5-anchor's comprehensive synonym coverage.

4 / 5

Distinctiveness Conflict Risk

A clear niche (statistical validation of variant predictors against DMS data) with domain-specific triggers, but the broad 'or any per-variant numeric score' clause and closely related sibling skills in the same collection (e.g. protein-sae-variant-interpretation) leave minor overlap risk, fitting the 4-anchor 'mostly distinct'.

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
mims-harvard/ToolUniverse
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.