CtrlK
BlogDocsLog inGet started
Tessl Logo

scientific-critical-thinking

Evaluate research rigor. Assess methodology, experimental design, statistical validity, biases, confounding, evidence quality (GRADE, Cochrane ROB), for critical analysis of scientific claims.

50

Quality

56%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./scientific-skills/Evidence Insight/scientific-critical-thinking/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

46%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-structured with real, clearly linked reference files, but it is markedly verbose—inlining definitions of concepts Claude already knows and an unrelated schematics section—and offers checklists of questions rather than a concrete, executable evaluation workflow.

Suggestions

Remove the inlined bias/fallacy/statistics definitions (concepts Claude already knows) and rely on the existing reference files, keeping only the decision-relevant checklist prompts in SKILL.md.

Delete or relocate the 'Visual Enhancement with Scientific Schematics' section; it is off-scope for critical thinking and cites a scripts/generate_schematic.py that is not present in the bundle.

Convert the claim-evaluation and critique sections from lists of questions into an explicit sequenced workflow with decision checkpoints (e.g., classify claim strength → match to evidence → flag red flags → emit structured critique).

DimensionReasoningScore

Conciseness

The ~570-line body extensively inlines concepts Claude already knows (defining every bias, fallacy, and p-value basics) and includes an off-topic 'Visual Enhancement with Scientific Schematics' section that references a non-existent scripts/generate_schematic.py, producing several padded, unnecessary sections.

2 / 5

Actionability

Provides a concrete five-part critique structure and per-domain checklists of questions to ask, but guidance is mostly prompts-to-consider rather than executable decision procedure, leaving key how-to-decide detail incomplete.

3 / 5

Workflow Clarity

Numbered sequences exist (claim evaluation, design process, critique structure) but they read as taxonomies; no explicit validation checkpoints or error-recovery feedback loops are present, though the task is non-destructive so the hard cap does not bind.

3 / 5

Progressive Disclosure

Well-sectioned overview with six clearly signaled, one-level-deep reference files that all exist on disk; the main gap is that large bias/fallacy/statistics taxonomies are inlined in the body duplicating the reference files rather than living only in references.

4 / 5

Total

12

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and well-targeted to a scientific-critique niche with named frameworks, but it lacks an explicit 'Use when...' trigger clause, leaving the activation guidance only weakly implied and capping completeness at 3.

Suggestions

Add an explicit 'Use when...' trigger clause naming concrete user situations, e.g. 'Use when reviewing research papers, assessing study rigor, or evaluating scientific claims and media reports of research.'

Include natural synonyms users might say (e.g. 'study quality', 'risk of bias', 'peer review', 'critical appraisal') to broaden trigger coverage.

Tighten the verb repetition ('Evaluate... Assess...') by varying the concrete actions to read more like distinct capabilities.

DimensionReasoningScore

Specificity

Lists several specific assessment domains (methodology, experimental design, statistical validity, biases, confounding, evidence quality) plus named frameworks (GRADE, Cochrane ROB); not quite comprehensive since the verb form is uniformly 'assess/evaluate'.

4 / 5

Completeness

Has a clear 'what' (evaluate rigor / assess the listed domains) but the 'when' is only weakly implied via 'for critical analysis of scientific claims' with no explicit 'Use when...' trigger clause, which caps completeness at 3.

3 / 5

Trigger Term Quality

Good coverage of natural terms a research-focused user would say ('research rigor', 'methodology', 'biases', 'evidence quality'), though synonyms and common variations are limited.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (scientific research critique anchored to GRADE/Cochrane ROB) with distinct triggers; minor overlap risk with general analysis skills keeps it just below 5.

4 / 5

Total

15

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (578 lines); consider splitting into references/ and linking

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 1 missing

Warning

Total

13

/

16

Passed

Repository
aipoch/medical-research-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.