CtrlK
BlogDocsLog inGet started
Tessl Logo

study-design-scale-selector

Determines the appropriate Risk of Bias assessment scale for a medical study based on its design (RCT, Cohort, etc.), using PubMed metadata lookup or text analysis. Use when the user wants to know which quality assessment tool to use for a specific paper (given PMID or abstract).

60

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./scientific-skills/Data Analysis/study-design-scale-selector/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

60%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The core skill is genuinely good — a clear conditional workflow from PMID lookup to text-analysis fallback to scale selection, backed by real, correctly split bundle files. It is dragged down by heavy auto-generated boilerplate (~100 of ~145 lines add nothing Claude does not already know) and several template artifacts: a truncated helper section, a reference to a nonexistent CONFIG block, a wrong-direction cross-reference, and a machine-specific example path.

Suggestions

Delete the generic template sections (Key Features, Implementation Details, Required Inputs, Output Contract, Validation and Safety Rules, Failure Handling, and the redundant When to Use/When Not to Use) and keep only the Workflow, Helper Scripts, and Quick Validation sections — this alone would move conciseness toward the top anchors.

Complete the truncated PDF Text Extraction section with the actual usage command from extract_pdf.py's docstring (`python scripts/extract_pdf.py <input PDF file> [--output <output file>]`), and remove the reference to the nonexistent in-file CONFIG block in the Example run plan.

Fix the misleading cross-reference 'See ## Workflow above for related details' (the Workflow section is below it) and replace the machine-specific example path (20260316/scientific-skills/...) with a path relative to the skill root.

DimensionReasoningScore

Conciseness

Roughly two-thirds of the body is template boilerplate that assumes no intelligence and pads the token budget: "Use this skill when the request matches its documented task boundary", "Packaged executable path(s): scripts/extract_pdf.py plus 1 additional script(s)", "Output discipline: keep results reproducible... avoid undocumented side effects", plus duplicated When to Use / When Not to Use / Required Inputs / Output Contract / Validation and Safety Rules / Failure Handling sections. This clearly matches the 'noticeably verbose; several unnecessary... padded sections' anchor; it is short of level 1 only because it does not explain basic domain concepts (what a PDF or an RCT is), it just restates generic policy.

2 / 5

Actionability

The core workflow is executable: `python scripts/selector.py "<PMID>"` is a copy-paste-ready command, the fallback gives concrete keywords ("Randomized controlled trial", "RCT", ...), the output is a literal JSON block, and the referenced bundle files (references/scale_rules.md, both scripts) all exist. It stops short of level 5 because the PDF Text Extraction section is truncated mid-sentence ("use extract_pdf.py to extract the text content before assessment:" with nothing following) and the Example run plan tells the user to edit an in-file CONFIG block that does not exist in either script.

4 / 5

Workflow Clarity

The four-step workflow is clearly sequenced with an explicit conditional checkpoint (if selector.py returns non-empty JSON, skip to step 3; otherwise fall back to text analysis), which is a genuine validation/branching step. It is not level 5 because there is no error-recovery loop for the PubMed failure path beyond falling back, the 'See ## Workflow above' cross-reference in Implementation Details points the wrong direction (Workflow appears below it), and the truncated helper section and phantom CONFIG reference add confusion.

4 / 5

Progressive Disclosure

Scored against the actual bundle: SKILL.md is an overview, the scale-selection rules are properly split into a real, well-signaled one-level-deep reference ([scale_rules.md](references/scale_rules.md)), and the executable logic lives in the two real scripts. This fits the 'good structure; most content is appropriately placed; references mostly clear; minor organization gaps' anchor rather than level 5, because the truncated PDF-extraction section and the ten generic boilerplate sections dilute navigation and bury the real workflow.

4 / 5

Total

14

/

20

Passed

Description

82%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that clearly and explicitly states both what the skill does and when to use it, with credible natural trigger terms and a distinct niche. The main weakness is limited specificity of the action list — it describes the outcome but not the fuller range of capabilities (design identification, PDF input, specific tool names).

Suggestions

Mention the study-design identification step as its own action (e.g., 'Identifies the study design from PubMed metadata or abstract text, then selects...') to raise action specificity.

Name the tool families the selector can return (e.g., 'RoB 2, NOS, ROBINS-I, QUADAS-2') so users searching by tool name also match.

Add 'study design' and 'risk of bias' synonyms such as 'RoB' or 'quality appraisal' to broaden natural keyword coverage.

DimensionReasoningScore

Specificity

The description names the domain and 1-2 concrete actions — "Determines the appropriate Risk of Bias assessment scale... using PubMed metadata lookup or text analysis" — but coverage is not comprehensive: it never mentions identifying the study design itself, PDF handling, or the tool families (RoB 2, NOS, ROBINS-I) it selects among. This matches the 'names domain and 1-2 concrete actions' anchor, and falls short of the level-4 anchor, which expects several listed specific actions.

3 / 5

Completeness

Both questions are explicitly answered: the 'what' is "Determines the appropriate Risk of Bias assessment scale for a medical study based on its design (RCT, Cohort, etc.)", and the 'when' is an explicit "Use when the user wants to know which quality assessment tool to use for a specific paper (given PMID or abstract)" with concrete trigger conditions. This matches the level-5 anchor; level 4 would require the 'when' to be less explicit, which it is not.

5 / 5

Trigger Term Quality

Natural user phrases are present — "which quality assessment tool", "Risk of Bias assessment scale", "RCT, Cohort", "PMID or abstract" — which a researcher would plausibly say. A few common variations are missing ("RoB", "Cochrane", "Newcastle-Ottawa", "study appraisal"), which keeps it at the 'good keyword coverage; a few natural terms missing' anchor rather than level 5's comprehensive synonym coverage.

4 / 5

Distinctiveness Conflict Risk

The description occupies a clear niche — selecting a risk-of-bias instrument for a specific medical paper — with distinct triggers (PMID, abstract, quality assessment tool) that are unlikely to fire for unrelated skills. It does not trigger for generic literature-search or PDF-processing skills, fitting the 'clear niche with distinct triggers; minimal conflict risk' anchor better than level 4's 'minor overlap risk'.

5 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
aipoch/medical-research-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.