CtrlK
BlogDocsLog inGet started
Tessl Logo

usmle-case-generator

Generate USMLE Step 1/2 style clinical cases with patient history, physical.

45

Quality

56%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./scientific-skills/Academic Writing/usmle-case-generator/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

42%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The skill has a strong executable core — accurate CLI examples, a parameters table, a detailed case structure, and a realistic example output — buried under heavy generic boilerplate and self-referential filler. Dangling references and a hardcoded non-portable path undermine trust in the concrete instructions.

Suggestions

Cut the generic template sections (Risk Assessment, Security Checklist, Evaluation Criteria, Lifecycle Status, Response Template, Output Requirements) that carry no USMLE-specific information, and consolidate the duplicated Usage/Example Usage and Quick Check/Audit-Ready Commands pairs into single sections.

Fix broken path references: remove the hardcoded 'cd "20260318/scientific-skills/..."' line, point 'pip install -r' at references/requirements.txt, and either create references/conditions/ or drop it from the References section.

Surface the unused bundle files (guidelines.md, sample_input.json, sample_output.json) with one-line pointers, and replace the inline Topics Covered list with a pointer to references/topics.json.

DimensionReasoningScore

Conciseness

Roughly half the body is generic template padding unrelated to USMLE case generation: 'Use this skill for academic writing tasks that require explicit assumptions, bounded scope, and a reproducible output format', the Risk Assessment table, Security Checklist, Evaluation Criteria, Lifecycle Status ('Next Review Date: 2026-03-06'), and a seven-part Response Template, plus duplicated sections (Quick Check vs Audit-Ready Commands are identical; Usage vs Example Usage overlap). This matches anchor 2 ('noticeably verbose; several unnecessary explanations or padded sections') — not 1 because there is a real, specific skill core (parameters, case structure, example output) rather than tutorializing concepts Claude already knows.

2 / 5

Actionability

There is genuinely concrete guidance — 'python scripts/main.py --step 1 --topic cardiology --difficulty medium' (verified against the actual script's argparse flags), a full Parameters table, and a worked Example Output — but key details are broken: the Example Usage 'cd "20260318/scientific-skills/..."' path is a non-portable hardcoded path, 'references/conditions/' does not exist in the bundle, and 'pip install -r requirements.txt' omits that the file lives under references/. This sits between anchor 3 (concrete guidance but incomplete, missing key details) and anchor 4; the dangling paths and duplicated abstract sections pull it to 3.

3 / 5

Workflow Clarity

The five-step Workflow is sequenced and there is a fallback path ('If execution fails or inputs are incomplete, switch to the fallback path...'), but the steps are abstract directives ('Validate that the request matches the documented scope') with no operational checkpoints, and the only validation is the py_compile smoke test — no step verifies the generated case output. This matches anchor 3 ('steps listed but validation gaps; checkpoints missing or implicit') — not 4 because the checkpoints are policy language rather than executable commands embedded in the sequence.

3 / 5

Progressive Disclosure

References are one level deep and mostly real (references/topics.json, references/case_templates.json, references/usmle_patterns.md all exist), but 'references/conditions/' is a dangling path, and three actual bundle files (guidelines.md, sample_input.json, sample_output.json) are never mentioned. Meanwhile content that belongs in those files is inlined — the 14-item Topics Covered list duplicates topics.json, and the security/risk boilerplate should live elsewhere or be cut. Anchor 3 ('some structure but could be better organized; references present but not clearly signaled; content that should be separate is inline') is the best fit.

3 / 5

Total

11

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description names a specific domain with recognizable trigger terms, but it is incomplete: it lacks any 'use when' guidance, appears truncated mid-sentence, and misses common synonyms like 'board exam' or 'practice questions'. It is functional but below the standard of the good examples.

Suggestions

Add an explicit trigger clause, e.g. 'Use when the user asks for USMLE practice questions, board exam preparation, or clinical vignettes.'

Complete the truncated capability list ('patient history, physical') to cover the documented output: physical exam, labs, imaging, multiple choice questions with answer explanations.

Include common user phrasings such as 'board exam prep', 'Step 2 CK', and 'clinical vignettes' to broaden natural trigger coverage.

DimensionReasoningScore

Specificity

The description names the domain ('USMLE Step 1/2 style clinical cases') and 1-2 concrete elements ('patient history, physical'), but the list appears truncated mid-sentence ('...patient history, physical.') and omits documented capabilities like labs, imaging, multiple choice questions, and answer explanations. This matches anchor 3 ('names domain and 1-2 concrete actions, but not comprehensive') — not 4 because it does not list several specific actions with only minor gaps, and not 2 because the domain and concrete actions are genuinely named.

3 / 5

Completeness

The 'what' is clear ('Generate USMLE Step 1/2 style clinical cases with patient history, physical.') but there is no 'when to use' clause anywhere — the judging guideline caps completeness at 3 for a missing 'Use when...' trigger. It is not 2 because the 'what' is specific rather than vague, and not 4 because the 'when' is entirely absent rather than merely under-specified.

3 / 5

Trigger Term Quality

Natural keywords like 'USMLE', 'Step 1/2', 'clinical cases', and 'patient history' are present and would be said by a target user, but common variations are missing: 'board exam', 'practice questions', 'vignettes', 'Step 2 CK', 'medical education'. Anchor 3 ('some relevant keywords but missing common variations or synonyms') fits; not 4 because several natural trigger phrases users would actually say are absent.

3 / 5

Distinctiveness Conflict Risk

'USMLE Step 1/2' and 'clinical cases' are niche-specific triggers unlikely to fire for unrelated skills, with only minor overlap risk against general medical-education or case-writing skills — matching anchor 4. Not 5 because without any 'use when' phrasing, broader requests like 'write a patient case' or 'medical practice questions' could plausibly match this or competing skills.

4 / 5

Total

13

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 1 missing

Warning

Total

14

/

16

Passed

Repository
aipoch/medical-research-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.