CtrlK
BlogDocsLog inGet started
Tessl Logo

result-reliability-checker

Assesses whether study results are trustworthy by auditing design integrity, sample structure, statistical handling, bias control, validation chain, and claim discipline. It identifies where results are robust, fragile, overfit, under-validated, or overclaimed. Always separate reported findings from reliability judgment. Never fabricate references, PMIDs, DOIs, trial identifiers, study features, or validation claims.

61

Quality

73%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./awesome-med-research-skills/Evidence Insight/result-reliability-checker/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Well-structured audit skill with clear sequencing, a self-check checkpoint, and excellent progressive disclosure via real reference modules, though it is somewhat verbose due to repeated prohibition lists.

Suggestions

Consolidate the repeated 'should not' / prohibition lists (Core Function, Hard Rules, What This Skill Should Not Do) into a single authoritative list to reduce redundancy.

Tighten the 'What This Skill Should Not Do' section since it largely restates the Hard Rules.

DimensionReasoningScore

Conciseness

The body is mostly directive and assumes Claude's knowledge (no explaining of p-values or AUROC), but it repeats the same prohibitions across 'Core Function should not', 'Hard Rules', and 'What This Skill Should Not Do', so it could be tightened.

3 / 5

Actionability

As an instruction-only skill it gives concrete, specific guidance — an 8-step procedure with explicit per-step checklists and a fully specified mandatory output structure (sections A–I).

4 / 5

Workflow Clarity

The 8 steps are explicitly ordered ('always run in order') with a self-critical final check (Step 8) as a validation checkpoint, though it lacks a true error-recovery feedback loop.

4 / 5

Progressive Disclosure

The body is a clear overview pointing to seven one-level-deep reference modules, each signalled with a clear 'Use for/when' purpose, and all referenced files exist in references/.

5 / 5

Total

16

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific, third-person, and well-differentiated, but it omits an explicit 'when to use' trigger clause, which caps its completeness score.

Suggestions

Add an explicit 'Use when...' clause naming natural user triggers (e.g., 'Use when checking whether a paper's results are reliable, robust, or overstated').

Include common user phrasings and synonyms like 'are these results trustworthy', 'is this study reliable', or 'does this paper overclaim'.

DimensionReasoningScore

Specificity

Lists multiple concrete audit actions — 'auditing design integrity, sample structure, statistical handling, bias control, validation chain, and claim discipline' plus 'identifies where results are robust, fragile, overfit, under-validated, or overclaimed' — giving comprehensive coverage.

5 / 5

Completeness

The 'what' is clear and detailed, but there is no explicit 'Use when...' or equivalent trigger guidance; per the rubric, a missing trigger clause caps completeness at 3.

3 / 5

Trigger Term Quality

Contains natural terms a user might say ('trustworthy', 'robust', 'fragile', 'overclaimed') but lacks an explicit 'Use when...' clause with synonyms and common user phrasings, so a few natural terms are missing.

4 / 5

Distinctiveness Conflict Risk

Targets a clear niche — medical research result-trustworthiness auditing — with distinct triggers and minimal overlap risk against generic skills.

5 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
aipoch/medical-research-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.