CtrlK
BlogDocsLog inGet started
Tessl Logo

tooluniverse-statistical-modeling

Statistical modeling — linear/logistic/ordinal/Poisson regression, ANOVA, Kruskal-Wallis, chi-square, Mann-Whitney, Cox survival, spline fits (R `ns()`), odds ratios, Cohen's d, F-statistic, p-value computation. Specializes in clinical-trial AE analysis (SDTM DM/AE), severity ordinal regression, and per-feature stat workflows.

58

Quality

67%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./plugins/tooluniverse/skills/tooluniverse-statistical-modeling/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exceptionally actionable skill body with executable commands, wrong/right examples, validation checkpoints, and a final checklist — its greatest strength. It is dragged down by heavy duplication (AE, spline, and ANOVA guidance each appear 2-4 times), generic statistics explanations Claude already knows, and references to six bundle files that do not exist.

Suggestions

Deduplicate: keep one authoritative section each for AE-severity analysis, spline fitting, and expression ANOVA (the PRIMARY SCRIPTS table plus one worked invocation), and delete the repeated copies in 'CRITICAL — Read before writing any code' and 'Analysis conventions'.

Fix the broken references: either add the missing anova_and_tests.md, common_patterns_summary.md, TOOLS_REFERENCE.md, QUICK_START.md, EXAMPLES.md, and test_skill.py to the bundle, or remove them from the File Structure map and body citations.

Move the long 'Analysis conventions' details and the Pearson four-variant protocol into a reference file, and trim textbook interpretation guidance (OR/HR meaning, AIC/BIC basics) that Claude already knows; also document or remove the unlisted sdtm_ordinal_logistic.py script.

DimensionReasoningScore

Conciseness

The ~660-line body duplicates whole topics: AE-severity guidance appears nearly verbatim in "Clinical trial AE analysis" and again in "Analysis conventions"; spline guidance is repeated across "PRIMARY SCRIPTS", "CRITICAL" item 3, "Bundled Scripts", and the strain co-culture section; the expression-ANOVA rules and their script invocations are likewise given twice. It also explains concepts Claude already knows ("OR = 1.0 means no association", "HR = 2.0 means twice the instantaneous risk", "Lower [AIC] is better"). This is noticeably verbose with several padded/duplicated sections — the 2 anchor — rather than 3, because the duplication alone would cut the file by roughly a third without losing information.

2 / 5

Actionability

Nearly every guidance block is copy-paste executable: full CLI invocations with flags for each bundled script (e.g. the logistic_regression_or.py call with --outcome-order, --encode-map, --interaction), the tu run command, executable Python snippets, and paired ❌ WRONG / ✅ RIGHT examples covering the common failure modes. This matches the fully-executable anchor; not 4 because there are no material gaps — even fallbacks (stat_tests.py for no-scipy environments) are covered.

5 / 5

Workflow Clarity

There is a clear sequence (RULE ZERO pre-computed-results check → primary scripts → Phase 0 data validation → Phase 1 model fitting → Phase 2 diagnostics → Phase 3 interpretation) with genuine checkpoints: Phase 0 outcome-variable verification, sanity heuristics that act as feedback loops ("F > 50 ... means you aggregated"), and a final Completeness Checklist. It falls short of 5 because the workflow is split across four competing entry-point sections (RULE ZERO, PRIMARY SCRIPTS, CRITICAL, Workflow) with overlapping instructions, leaving some ambiguity about which path governs, and the diagnostics phase defers detail to a missing reference file.

4 / 5

Progressive Disclosure

The six references/*.md files cited in the body do exist and are one level deep, but the body also cites files absent from the bundle — anova_and_tests.md, common_patterns_summary.md, TOOLS_REFERENCE.md, QUICK_START.md, EXAMPLES.md, and test_skill.py (all listed in the File Structure map) — so several navigation pointers dead-end. Meanwhile ~200 lines of repeated analysis conventions remain inlined in SKILL.md where they belong in a reference file. This fits the 3 anchor (structure present, but content that should be separate is inline and the reference set is partly broken); not 4 because of the missing-file references, and not 2 because the existing references are clearly signaled and the file is well sectioned.

3 / 5

Total

14

/

20

Passed

Description

71%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A dense, concrete capability inventory with excellent method-level specificity and a well-defined clinical-trial niche. Its main defect is the missing explicit 'Use when...' trigger clause, and a few common user phrasings (hazard ratio, t-test, power/sample size) are absent.

Suggestions

Append an explicit trigger clause, e.g. "Use when the user asks for regression, odds ratios, hazard ratios, p-values, ANOVA, survival analysis, or statistical testing of biomedical/clinical-trial data."

Add commonly-used natural synonyms currently missing: "hazard ratio", "Kaplan-Meier", "t-test", "sample size / power analysis".

DimensionReasoningScore

Specificity

The description enumerates many concrete capabilities — "linear/logistic/ordinal/Poisson regression, ANOVA, Kruskal-Wallis, chi-square, Mann-Whitney, Cox survival, spline fits (R ns()), odds ratios, Cohen's d, F-statistic, p-value computation" — plus a concrete specialization ("clinical-trial AE analysis (SDTM DM/AE), severity ordinal regression"). This matches the comprehensive-coverage anchor; it is not a 4 because there are no meaningful gaps in the method list relative to the skill's actual scope.

5 / 5

Completeness

The "what" is clearly answered (a full method inventory), but there is no "Use when..." clause or equivalent explicit trigger guidance — "Specializes in clinical-trial AE analysis" describes a focus area, not when to invoke the skill. Per the judging guideline, a missing 'Use when' clause caps completeness at 3; it is above 2 because the what-half is explicit and detailed.

3 / 5

Trigger Term Quality

Strong natural terms a user would actually say ("odds ratio", "p-value", "ANOVA", "chi-square", "Cox survival", "Cohen's d", "spline"), but common phrasings like "hazard ratio", "t-test", "sample size / power analysis", and "Kaplan-Meier" are absent. Good coverage with a few natural terms missing fits the 4 anchor; not 5 because those frequently-used synonyms are missing, not 3 because the included keywords are varied and user-facing rather than generic.

4 / 5

Distinctiveness Conflict Risk

The SDTM DM/AE, severity-ordinal-regression, and R-ns() spline niche makes it mostly distinct from generic data-analysis or bioinformatics skills. Minor overlap risk remains with any general statistics skill (ANOVA, regression, p-value are broad triggers); not 5 because those generic method names would also plausibly fire a rival statistics skill.

4 / 5

Total

16

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (660 lines); consider splitting into references/ and linking

Warning

Total

15

/

16

Passed

Repository
mims-harvard/ToolUniverse
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.