CtrlK
BlogDocsLog inGet started
Tessl Logo

statistical-analysis-advisor

Recommends appropriate statistical methods (T-test vs ANOVA, etc.) based.

49

Quality

62%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./scientific-skills/Data Analysis/statistical-analysis-advisor/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

53%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body has a solid core — a verified executable script, accurate API examples, and well-organized one-level-deep references — but it is buried in duplicated commands, circular self-references, dated machine-generated paths, and generic boilerplate sections. The workflow is sequenced but its checkpoints are abstract rather than concrete.

Suggestions

Cut the boilerplate sections (Risk Assessment, Security Checklist, Lifecycle Status, Evaluation Criteria, Response Template) and the circular 'See ## X above' cross-references, and deduplicate the py_compile command to a single occurrence.

Fix the command examples: remove the bogus `--help` flag (the script has no argparse and ignores it), document that `python scripts/main.py` runs a canned demo, and delete the dated `cd "20260318/..."` path.

Ground the Workflow steps in concrete validation, e.g. 'run `python -m py_compile scripts/main.py`, then execute the advisor with the confirmed parameters and verify the recommendation's assumptions section before returning it'.

DimensionReasoningScore

Conciseness

The ~250-line body is noticeably padded: the py_compile command appears three times (Quick Check, Audit-Ready Commands, Example Usage), the frontmatter description is repeated verbatim in 'When to Use', several sections contain only circular cross-references ("See `## Prerequisites` above for related details"), and boilerplate sections (Risk Assessment, Security Checklist, Lifecycle Status, Evaluation Criteria, Response Template) add no task-specific value. This matches 'Noticeably verbose; several unnecessary explanations or padded sections'; it is not 1 because genuinely useful content (Capabilities, Usage code, Input Parameters, References) is interleaved with the padding.

2 / 5

Actionability

The Usage section shows copy-paste-ready code (StatisticalAdvisor.recommend_test / check_assumptions / calculate_power) that exactly matches the real API in scripts/main.py, plus a concrete parameter table and runnable verification commands — fitting 'Mostly executable guidance; concrete code or commands with minor gaps'. It is not 5 because `python scripts/main.py --help` is misleading (the script has no argparse, so --help is silently ignored and the demo runs instead), the bare `python scripts/main.py` runs a canned demo rather than the documented input-driven run plan, and the example `cd "20260318/scientific-skills/..."` path is a wrong machine-generated artifact.

4 / 5

Workflow Clarity

The Workflow section lists a coherent 5-step sequence with stop-early and fallback rules, matching 'Steps listed but validation gaps; sequence present but checkpoints missing or implicit'. It is not 4 because the checkpoints are abstract policy statements ('validate that the request matches the documented scope') rather than concrete verify commands tied to steps, and the triplicated compile-check sections blur the actual entry path; it is above 2 because a real, ordered sequence with explicit error-handling behavior exists.

3 / 5

Progressive Disclosure

The bundle structure is sound: SKILL.md acts as the overview, detailed material lives in three real, substantive, one-level-deep files (references/statistical_tests_guide.md, assumption_tests.md, power_analysis_guide.md — each clearly signaled under References), and executable logic is separated into scripts/main.py. This fits 'Good structure; most content appropriately placed; references mostly clear; minor organization gaps'. It is not 5 because the body also inlines substantial boilerplate (risk, security, lifecycle, evaluation sections) that adds navigation noise rather than earning its place in SKILL.md.

4 / 5

Total

13

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description identifies a real niche with a couple of concrete test names, but it is truncated mid-sentence, lists only one capability, and entirely lacks 'when to use' trigger guidance. It reads as an incomplete draft rather than an effective routing description.

Suggestions

Finish the truncated sentence: 'Recommends appropriate statistical methods (T-test vs ANOVA, etc.) based on data type, distribution, and study design.'

Add a 'Use when...' clause with natural triggers, e.g. 'Use when deciding between statistical tests, checking test assumptions, or determining sample size/power.'

Include the other documented capabilities (assumption checking, power/sample-size analysis) so the 'what' is comprehensive rather than a single action.

DimensionReasoningScore

Specificity

The description names the domain and a single action — "Recommends appropriate statistical methods (T-test vs ANOVA, etc.)" — with two concrete test names, matching the anchor for domain plus 1-2 concrete actions but not comprehensive. It is not a 4 because only one action verb appears ("recommends"), with no coverage of the skill's other capabilities (assumption checking, power analysis); not a 2 because naming T-test and ANOVA is more concrete than the generic 'Names the domain but actions are minimal' example.

3 / 5

Completeness

A clear "what" is present (recommends statistical methods), but there is no "Use when..." clause or equivalent trigger guidance, which caps completeness at 3 per the judging guidelines. It is also truncated mid-sentence ("...etc.) based."), which weakens even the "what" side, ruling out 4; the clear domain statement keeps it above 2, where 'when' would be absent AND 'what' vague.

3 / 5

Trigger Term Quality

"statistical methods", "T-test", and "ANOVA" are terms users would naturally say, but common variations are missing — no "statistical test", "significance", "compare groups", "p-value", or "power analysis". This matches 'Some relevant keywords but missing common variations or synonyms'; it falls short of 4 because the keyword set is only two test names plus a generic noun, with no synonyms for the task itself.

3 / 5

Distinctiveness Conflict Risk

Statistical test recommendation is a distinct niche and the specific test names (T-test, ANOVA) separate it from generic data-analysis or document skills, fitting 'Mostly distinct; minor overlap risk with closely related skills'. It is not 5 because the trigger vocabulary is narrow (two tests) with no scope language distinguishing it from broader statistics or data-analysis skills.

4 / 5

Total

13

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
aipoch/medical-research-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.