CtrlK
BlogDocsLog inGet started
Tessl Logo

evidence-level-ranker

Ranks papers by evidence family, methodological quality tier, validation depth, and claim discipline; assigns anchor, context-setting, mechanistic support, or caution citation roles; prevents prestige-based or design-label-based ranking errors.

60

Quality

71%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./awesome-med-research-skills/Evidence Insight/evidence-level-ranker/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, actionable instruction skill with a clear sequenced workflow and mandatory output format, weakened mainly by repetitive restatement across the Execution, Hard Rules, 'What This Skill Should Not Do', and 'Quality Standard' sections.

Suggestions

Consolidate the redundant prohibitions: fold the 17 Hard Rules and 'What This Skill Should Not Do' bullets into the relevant Execution steps rather than restating them in three separate sections.

Add one short worked example (a small mixed paper set with a sample A–J ranked output) to lift actionability from concrete guidance to fully exemplified.

Move the detailed rule material into the existing reference files (which are currently thin stubs) so the SKILL.md body stays a lean overview and the split is better balanced.

DimensionReasoningScore

Conciseness

The body is mostly on-topic and avoids explaining basics Claude already knows, but it is padded by substantial redundancy: the 17 Hard Rules restate the Execution steps, 'What This Skill Should Not Do' restates the Hard Rules, and 'Quality Standard' restates the Core Function.

3 / 5

Actionability

Concrete and actionable guidance throughout — an 8-step workflow, specific quality dimensions to check, named overclaim patterns, defined citation roles, a mandatory A–J output structure, and a copy-paste out-of-scope response template — with the minor gap of no worked example showing a sample ranked output.

4 / 5

Workflow Clarity

A clearly sequenced 8-step process with checklists (Step 3 dimensions, overclaim patterns) and verification guards (input validation, 'unverified' labeling, limitations step); it lacks an explicit validate-fix-retry feedback loop, though the task is analytical rather than destructive.

4 / 5

Progressive Disclosure

All nine referenced files exist and are well-signaled with one-level-deep navigation mapped to specific steps, but the body inlines heavy material (8 detailed steps, 17 hard rules) while the reference files are thin (9-17 lines), so the content split could be better balanced.

4 / 5

Total

15

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A specific, well-scoped description with concrete actions and a distinct niche, but it lacks any explicit 'when to use' trigger guidance and relies on somewhat technical terminology over natural user phrasing.

Suggestions

Add an explicit 'Use when...' clause, e.g. 'Use when ranking a set of papers by evidence level or deciding which to cite first for a manuscript, review, or protocol.'

Weave in natural trigger phrases users actually say, such as 'evidence level', 'evidence strength', 'citation priority', and 'literature shortlist'.

Lead with the highest-intent user phrasing before the technical dimensions so the description reads as a trigger first.

DimensionReasoningScore

Specificity

Lists multiple concrete actions across the full ranking task ('Ranks papers by evidence family, methodological quality tier, validation depth, and claim discipline; assigns anchor, context-setting, mechanistic support, or caution citation roles; prevents prestige-based or design-label-based ranking errors'), giving comprehensive coverage of what the skill does.

5 / 5

Completeness

The 'what' is clearly and concretely stated, but there is no 'Use when...' clause or equivalent explicit trigger guidance, so per the rubric completeness is capped at 3.

3 / 5

Trigger Term Quality

Good keyword coverage with domain-relevant terms ('papers', 'evidence family', 'citation roles', 'ranking'), but it leans technical and omits common natural phrases users say such as 'evidence level', 'evidence strength', 'citation priority', or 'literature'.

4 / 5

Distinctiveness Conflict Risk

The niche is sharply defined (evidence-level ranking with anchor/context/mechanistic/caution citation roles) with distinct triggers and minimal overlap risk with adjacent literature or citation skills.

5 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
aipoch/medical-research-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.