CtrlK
BlogDocsLog inGet started
Tessl Logo

tooluniverse-statistical-modeling

Statistical modeling and regression analysis for biomedical data. Linear/logistic/ordinal regression, Cox proportional hazards, mixed-effects models, ANOVA, with odds ratios, hazard ratios, confidence intervals, and model diagnostics.

57

Quality

72%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./scientific-skills/Data Analysis/tooluniverse-statistical-modeling/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

70%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers strong, actionable statistical guidance with an exemplary phased workflow, validation checkpoints, and a completeness checklist. Its weaknesses are notable duplication (Quick Start vs Phase 1, ANOVA guidance vs Pattern 5, benchmark-specific padding) and a progressive-disclosure chain whose referenced files are missing from the bundle entirely.

Suggestions

Ship the referenced bundle files or remove the pointers: every reference (`references/*.md`, `scripts/*.py`, `QUICK_START.md`, `EXAMPLES.md`, `TOOLS_REFERENCE.md`) currently points to a nonexistent file.

De-duplicate: keep one canonical logistic/Cox example (drop the Quick Start copies of Phase 1 code) and merge Pattern 5 into the Phase 1 multi-feature ANOVA section instead of restating the decision tree and per-feature guidance twice.

Fix the two code defects: replace the broken `results = cph.check_assumptions(...)` / `len(results)` usage, and define or parameterize `expression_matrix` in the ANOVA Method A/B snippets; move benchmark-specific expected values (e.g. bix-36-q1, 0.76–0.78) into a reference file.

DimensionReasoningScore

Conciseness

Mostly efficient — no explanation of basic concepts and code is dense — but there is real duplication: the Quick Start logistic/Cox code is repeated nearly verbatim in Phase 1, the per-feature ANOVA guidance appears in full in Phase 1 and again as Pattern 5 ("default to per-feature ANOVA" stated three times), and benchmark-specific walkthroughs ("BixBench bix-36-q1... Expected: 0.76-0.78") read as over-fitted padding. Falls between anchor 2 (several padded sections) and anchor 4 (minor trimming); anchor 3 fits best.

3 / 5

Actionability

Concrete, mostly executable guidance for every model family (statsmodels formulas, lifelines Cox, scipy ANOVA). Not anchor 5: the ANOVA Method A/B code depends on an undefined `expression_matrix` variable, and `cph.check_assumptions()` returns None so `if len(results) == 0` would raise TypeError — minor but real gaps in copy-paste readiness.

4 / 5

Workflow Clarity

Clear phased sequence (Phase 0 data validation → Phase 1 model fitting → Phase 2 diagnostics → Phase 3 interpretation) with an explicit up-front validation checkpoint, a model-selection decision tree, error-recovery guidance for convergence/separation/collinearity, and an end-of-task completeness checklist — matches anchor 5 exactly.

5 / 5

Progressive Disclosure

References are clearly signaled one level deep and a file-structure map is provided, but scored against the actual bundle: none of the referenced files exist (`references/`, `scripts/`, `QUICK_START.md`, `EXAMPLES.md`, `TOOLS_REFERENCE.md` are all absent), so every pointer is broken and much of the content that the structure says is delegated is inlined at length in SKILL.md — anchor 3, not 4, because the disclosure chain does not resolve.

3 / 5

Total

15

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, specific description that names its domain and concrete statistical capabilities in third person, with good natural trigger terms. Its main weaknesses are the absence of an explicit 'Use when...' clause and a few missing natural synonyms (e.g., "survival analysis") that would push it to the top level.

Suggestions

Add an explicit trigger clause, e.g. "Use when the user asks for odds ratios, hazard ratios, p-values, or asks to fit regression, survival, or mixed-effects models on biomedical data."

Include the natural synonyms users commonly say, especially "survival analysis", "Kaplan-Meier", and basic tests (t-tests, chi-square), to broaden trigger coverage and close the capability gaps.

Clarify the boundary against prediction/ML use cases in the description (e.g. "for statistical inference, not ML prediction") to further reduce overlap with general data-analysis skills.

DimensionReasoningScore

Specificity

The description enumerates multiple concrete capabilities ("Linear/logistic/ordinal regression, Cox proportional hazards, mixed-effects models, ANOVA, with odds ratios, hazard ratios, confidence intervals") but omits capabilities the body advertises as core features (t-tests, chi-square, Kaplan-Meier), so coverage has minor gaps — anchor 4, not the comprehensive anchor 5.

4 / 5

Completeness

The 'what' is clear and concrete, but the description contains no 'Use when...' clause or equivalent explicit trigger guidance; per the judging guidelines this caps completeness at 3 even though the 'what' is well stated.

3 / 5

Trigger Term Quality

Includes strong natural phrases users would say ("regression analysis", "logistic", "Cox proportional hazards", "ANOVA", "odds ratios", "hazard ratios"), but misses common synonyms and variations such as "survival analysis", "fit a model", or "p-values" — good coverage with a few natural terms missing (anchor 4), not comprehensive (anchor 5).

4 / 5

Distinctiveness Conflict Risk

"biomedical data" plus named model families (Cox proportional hazards, mixed-effects, ANOVA) carves a mostly distinct niche with clear triggers; minor overlap risk remains with a general statistics or data-analysis skill — anchor 4, not the fully distinct anchor 5.

4 / 5

Total

15

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (578 lines); consider splitting into references/ and linking

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 8 missing

Warning

Total

13

/

16

Passed

Repository
aipoch/medical-research-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.