CtrlK
BlogDocsLog inGet started
Tessl Logo

claim-strength-calibrator

Calibrates manuscript claim strength so wording matches the actual evidence level, study design, and validation status.

51

Quality

56%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./awesome-med-research-skills/Academic Writing/claim-strength-calibrator/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

60%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-structured, actionable instruction-only skill with a clear sequenced workflow and a validation-first checkpoint, backed by real one-level-deep reference files. Its main weakness is substantial redundancy — the same evidence levels, overclaim patterns, and severity classes are restated across several sections.

Suggestions

Consolidate the duplicated enumerations: state evidence levels, overclaim patterns, and severity classes once and reference the corresponding file instead of restating them in 'Task', 'Important Distinctions', 'Execution', and 'Mandatory Output Structure'.

Merge or trim the largely overlapping 'Hard Rules', 'What This Skill Should Not Do', and 'Quality Standard' sections to reduce repeated content.

Move the inline evidence-level and overclaim-pattern lists entirely into their reference files, keeping the body as an overview that signals when to consult each reference.

DimensionReasoningScore

Conciseness

The same enumerations recur across multiple sections — evidence levels in 'Task', 'Important Distinctions', 'Execution Step 3', and 'Mandatory Output Structure'; overclaim patterns in 'Task', 'Reference Module Integration', 'Execution Step 4', and 'Mandatory Output Structure'; plus duplicated 'Hard Rules', 'What This Skill Should Not Do', and 'Quality Standard' — creating noticeably verbose, redundant padding that could be consolidated.

2 / 5

Actionability

Provides concrete instruction-only guidance: explicit 9-step execution, a mandatory 8-part output structure (sections A–H), defined severity buckets (major/moderate/minor/unclear), and concrete overclaim detection patterns, giving mostly executable guidance with only minor gaps.

4 / 5

Workflow Clarity

Sequences a clear 9-step process with an explicit validation checkpoint in 'Step 1 — Clarify before calibrating' ('do not immediately produce a full calibration review… ask focused follow-up questions'), plus a defined final output structure; missing only an explicit post-output error-recovery loop, so it sits just below the top anchor.

4 / 5

Progressive Disclosure

References seven one-level-deep files under 'Reference Module Integration', each clearly signaled with a purpose bullet, and all referenced paths verified to exist; however content that also lives in those references (evidence levels, overclaim patterns, severity classes) is inlined in the body, leaving a minor organization gap.

4 / 5

Total

14

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description states a clear, specific 'what' tied to evidence level, study design, and validation status, but it lacks any 'Use when…' trigger guidance and natural user phrasings. It is distinct within its niche yet under-specified on when to invoke it.

Suggestions

Add an explicit 'Use when…' clause naming natural user phrasings such as 'when the user asks whether claims are overstated, to calibrate discussion tone, or to check for causal/translational overclaiming'.

Expand the action list from the single verb 'Calibrates' to several concrete actions (e.g., 'identifies overclaim patterns, maps claims to evidence levels, proposes re-worded bounds, classifies severity').

Include common trigger synonyms a user would actually say ('overclaim', 'tone', 'reviewer criticism', 'causal inflation').

DimensionReasoningScore

Specificity

Quotes 'Calibrates manuscript claim strength so wording matches the actual evidence level, study design, and validation status' — names the domain plus concrete calibration dimensions, but offers only a single composite action ('Calibrates') rather than a list of specific actions, matching the 'Names domain and 1-2 concrete actions' anchor.

3 / 5

Completeness

Clearly answers 'what' (calibrate wording to match evidence/design/validation) but provides no 'Use when…' clause or equivalent trigger guidance, and the rubric caps completeness at 3 when that trigger guidance is missing.

3 / 5

Trigger Term Quality

Contains relevant domain keywords ('manuscript claim strength', 'evidence level', 'validation status') but omits the natural phrasings a user would actually say ('are our claims overstated', 'calibrate the tone', 'overclaiming'), so it has some relevant keywords while missing common variations.

3 / 5

Distinctiveness Conflict Risk

Targets a specific niche — manuscript claim-strength calibration for biomedical/academic writing — making it mostly distinct from generic editing skills with only minor overlap risk.

4 / 5

Total

13

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
aipoch/medical-research-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.