CtrlK
BlogDocsLog inGet started
Tessl Logo

tooluniverse-variant-predictor-dms-validation

Validate a variant-effect predictor (AlphaMissense, ESM-C SAE, ESM logits, EVE, conservation scores, or any per-variant numeric score) against experimental deep mutational scanning (DMS) data. Computes per-variant predictor scores, splits variants into neutral vs disruptive groups by DMS effect, runs a Mann-Whitney U test on the predictor scores, and sweeps the stratification thresholds for robustness. Use when you need to know whether a predictor's scores track real functional disruption on a specific protein.

71

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A high-quality, highly actionable skill body: complete executable code, explicit validation gates with error-recovery guidance for batch predictor computes, honest limitations, and clear sequencing. The main gaps are minor: a variable-name mismatch in the Step 6 plotting snippet, some trimmable narrative, and a large single-file body that could split predictor-specific details into reference files.

Suggestions

Fix the Step 6 snippet to use the defined variables (s_neutral_finite / s_disruptive_finite) so it is copy-paste executable.

Move the AlphaMissense response-shape documentation and bin-parsing code (Predictor option B) into a references/ file, keeping a short pointer plus the decision table in SKILL.md.

Trim narrative asides such as the motivating KRAS eval anecdote in Step 4 down to the one-line takeaway ('give both MWU and Spearman unless one is clearly inappropriate').

DimensionReasoningScore

Conciseness

The body is dense with genuinely non-obvious domain knowledge (tool signatures, Forge cost tables, API response shapes, sign conventions) and avoids explaining concepts Claude already knows. Not a 5 because there is some narrative padding that could be trimmed — e.g. the motivating KRAS eval anecdote in Step 4 ('The eval that motivated this skill... got qualitatively similar answers') and some justificatory prose around the AlphaMissense options — but it is efficient overall, well above the 'noticeably verbose' level 3.

4 / 5

Actionability

Nearly all guidance is executable: complete function definitions (categorize, sae_drop_per_variant, parse_bin_list, am_per_variant_matrix), concrete tool calls with real signatures and return shapes, and a runnable sweep. Not a 5 because Step 6's plotting snippet references undefined variables (s_neutral, s_disruptive, p — the defined names are s_neutral_finite/s_disruptive_finite), and option A never shows how wt_pooled/variant_pooled are obtained, so it is not fully copy-paste ready.

4 / 5

Workflow Clarity

Steps 1–6 are explicitly sequenced, and Step 3.5 is a hard validation gate for a batch compute ('Insufficient predictor scores... INSTRUCTS to INVESTIGATE THE PREDICTOR COMPUTATION STEP before continuing') with a feedback loop — 'do not paper over... debug the predictor computation step (re-run, check API keys, check log files)'. The mandatory sign-convention double-check and the robustness sweep add further checkpoints, matching the top anchor including the batch-operation validation requirement.

5 / 5

Progressive Disclosure

No bundle files exist, and the single SKILL.md is well structured with clear section headers, tables, and a cross-reference table pointing one level deep to sibling skills. Not a 5 because the file is ~440 lines of monolithic content — the per-predictor option details (especially the AlphaMissense response shape and bin-parsing code) would naturally live in a references/ file that the body currently inlines.

4 / 5

Total

17

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person voice, concrete multi-step capability statement, an explicit 'Use when' trigger clause, and named tools that make the niche unmistakable. The only weakness is slightly thin synonym coverage for the trigger phrasing (e.g., 'benchmark', 'correlate with DMS').

Suggestions

Add common user phrasings as trigger synonyms, e.g. 'Use when benchmarking or correlating a variant-effect predictor against DMS data'.

DimensionReasoningScore

Specificity

The description lists multiple concrete, specific actions — 'Computes per-variant predictor scores, splits variants into neutral vs disruptive groups by DMS effect, runs a Mann-Whitney U test on the predictor scores, and sweeps the stratification thresholds for robustness' — with comprehensive coverage and no vague padding. Not a 4 because the action list is complete rather than having minor gaps.

5 / 5

Completeness

Clearly answers both questions: 'what' via the enumerated concrete actions and 'when' via the explicit trigger clause 'Use when you need to know whether a predictor's scores track real functional disruption on a specific protein.' Both are concrete and explicit, matching the top anchor.

5 / 5

Trigger Term Quality

Good keyword coverage: 'variant-effect predictor', 'AlphaMissense', 'ESM-C SAE', 'deep mutational scanning (DMS)', 'Mann-Whitney U test', 'functional disruption' are natural terms a researcher would say. Not a 5 because common synonyms like 'benchmark', 'correlate with', or 'validate against DMS' phrasings users might naturally use are absent.

4 / 5

Distinctiveness Conflict Risk

The description carves a clear niche — validating per-variant numeric predictors against experimental DMS data — naming specific tools (AlphaMissense, ESM-C SAE, EVE), so it is unlikely to trigger for the wrong skill. Not a 4 because the tool names and statistical framing leave minimal overlap with adjacent analysis skills.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
mims-harvard/ToolUniverse
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.