CtrlK
BlogDocsLog inGet started
Tessl Logo

scientific-critical-thinking

Evaluate research rigor. Assess methodology, experimental design, statistical validity, biases, confounding, evidence quality (GRADE, Cochrane ROB), for critical analysis of scientific claims.

55

Quality

69%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./scientific-skills/Evidence Insight/scientific-critical-thinking/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

56%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-organized with genuinely actionable evaluation frameworks and correctly wired reference files, but it substantially violates progressive disclosure by inlining large bias/fallacy/GRADE catalogs that duplicate its own references and restate knowledge Claude already has. The schematic-generation section references a nonexistent script, undermining both actionability and navigation.

Suggestions

Cut Sections 2, 3, 4, and 5 down to short procedural summaries (what to check and when) and delegate the full bias, statistical-pitfall, GRADE, and fallacy catalogs to their existing reference files, which already contain that content.

Remove or fix the 'Visual Enhancement with Scientific Schematics' section: 'scripts/generate_schematic.py' does not exist in this bundle, so the command is not executable as written.

Add a single end-to-end evaluation workflow (read paper → select applicable checks → apply critique output structure → note uncertainty) so the capability sections feed one coherent process, and include one worked example of a critique.

DimensionReasoningScore

Conciseness

The ~570-line body inlines large catalogs that both duplicate the provided reference files and cover concepts Claude already knows: the bias taxonomy (Section 2), statistical pitfalls (Section 3), GRADE criteria (Section 4), and the logical fallacy catalog (Section 5) — roughly 200 lines that belong in references. This matches anchor 2 ('noticeably verbose; several unnecessary explanations or padded sections') rather than 3, since the padding is extensive rather than incidental.

2 / 5

Actionability

As an instruction-only skill it provides concrete, actionable guidance: numbered evaluation checklists, a structured critique output format, and a specific claim-evaluation process. It does not reach 5 because the only command in the body ('python scripts/generate_schematic.py') references a script that does not exist in the bundle, and the checklists would benefit from worked examples on a sample paper.

4 / 5

Workflow Clarity

Multi-step processes are clearly sequenced (claim evaluation steps, design process steps) and the 'When Providing Critique' section gives an explicit five-part output structure, with 'When Uncertain' covering contingency handling. Not 5 because there is no explicit end-to-end workflow tying capability selection to the critique output, and the schematic-generation workflow lacks a verification step for its broken script path.

4 / 5

Progressive Disclosure

All six referenced files (scientific_method.md, common_biases.md, statistical_pitfalls.md, evidence_hierarchy.md, logical_fallacies.md, experimental_design.md) exist, are one level deep, and are clearly signaled both per-section and in a Reference Materials section. However, significant content that should live in those references is inlined in the body, and the body cites 'scripts/generate_schematic.py' plus a 'scientific-schematics' skill that are not part of this bundle, which is a dangling navigation path.

3 / 5

Total

13

/

20

Passed

Description

71%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and framework-anchored, clearly communicating what the skill does with concrete domain vocabulary. Its main weakness is the absence of an explicit 'Use when...' trigger clause, which caps completeness and slightly limits discoverability for users who phrase requests in everyday terms.

Suggestions

Add an explicit 'Use when...' clause naming natural trigger phrases, e.g. 'Use when the user asks to critique, review, or evaluate a scientific study or research paper, assess risk of bias, or appraise evidence quality.'

Include everyday synonyms users actually say — 'research paper', 'study', 'critique this study', 'peer review', 'is this study reliable' — alongside the technical terms to broaden trigger coverage.

Tighten the broad opening phrase 'Evaluate research rigor' so the first clause is as concrete as the rest of the description.

DimensionReasoningScore

Specificity

The description lists multiple specific concrete actions ('Assess methodology, experimental design, statistical validity, biases, confounding, evidence quality (GRADE, Cochrane ROB)') with named frameworks, giving comprehensive coverage of the skill's capabilities. It does not fall to anchor 4, which requires minor gaps in coverage.

5 / 5

Completeness

The 'what' is clear and specific, but there is no 'Use when...' clause or equivalent explicit trigger guidance; the trailing phrase 'for critical analysis of scientific claims' only weakly implies when to use the skill. Per the judging guidelines, a missing explicit trigger clause caps completeness at 3.

3 / 5

Trigger Term Quality

Good keyword coverage including 'methodology', 'experimental design', 'statistical validity', 'biases', 'confounding', 'evidence quality', 'GRADE', and 'Cochrane ROB', but common natural phrasings users would say ('research paper', 'study', 'critique this paper', 'peer review') are missing. Not score 3 since the present terms are specific and domain-relevant rather than generic.

4 / 5

Distinctiveness Conflict Risk

Named frameworks (GRADE, Cochrane ROB) and terms like 'confounding' carve out a fairly distinct niche with minimal conflict risk against unrelated skills. It stays at 4 rather than 5 because the broad opening ('Evaluate research rigor') could overlap with general research-assistance or peer-review skills.

4 / 5

Total

16

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (578 lines); consider splitting into references/ and linking

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 1 missing

Warning

Total

13

/

16

Passed

Repository
aipoch/medical-research-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.