CtrlK
BlogDocsLog inGet started
Tessl Logo

scientific-critical-thinking

Evaluate scientific claims and evidence quality. Use for assessing experimental design validity, identifying biases and confounders, applying evidence grading frameworks (GRADE, Cochrane Risk of Bias), or teaching critical analysis. Best for understanding evidence quality, identifying flaws. For formal peer review writing use peer-review.

64

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./scientific_writer/.claude/skills/scientific-critical-thinking/SKILL.md

The canonical home for this skill is scientific-critical-thinking in K-Dense-AI/scientific-agent-skills

SKILL.md
Quality
Evals
Security

Quality

Content

61%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured overview body with exemplary reference bundling (all files exist, one level deep, well described) and concrete critique scaffolding. Main weaknesses are duplicated reference listings and trigger content, the absence of an explicit end-to-end workflow, and no worked example.

Suggestions

Collapse the 'Reference Materials' section (lines 136-148) into the 'Core Capabilities' section — the same six files are listed twice with overlapping descriptions; keep one annotated list.

Add a short end-to-end workflow before 'Application Guidelines' (e.g., read the study → grep the relevant reference by capability → classify concerns by severity → emit the 5-part critique), so the operating sequence is explicit rather than implied by output format.

Trim the 'When to Use This Skill' bullet list and the 'Remember' section, which repeat the frontmatter description and the 'Be Constructive' principles (e.g., 'Recognize that all research has limitations' appears twice).

DimensionReasoningScore

Conciseness

Procedural guidance (Application Guidelines, When Uncertain, critique structure) is efficient, but there is noticeable redundancy: the six reference files are listed and described twice — once under 'Core Capabilities' (lines 67-72) and again in a verbose 'Reference Materials' section (lines 136-148); 'When to Use This Skill' re-lists triggers already in the description; and 'Recognize that all research has limitations' appears in both 'Be Constructive' and 'Remember'. This sits between anchor 2 (several padded sections) and anchor 4 (minor trims), so 3 fits best.

3 / 5

Actionability

Concrete, executable guidance throughout: a 5-part critique output structure with severity tiers defined, exact phrasing templates for uncertainty ('This could be X or Y; additional information needed is Z'), a ready grep command ('grep -r "pattern" references/'), and a copy-paste-ready bash command for optional figures. It falls short of anchor 5 because there is no worked example applying the framework to an actual study, and much of the 'Remember' section is principle statements rather than instructions.

4 / 5

Workflow Clarity

The critique structure (Summary → Strengths → Concerns by severity → Recommendations → Overall Assessment) is a clear sequence, but there is no end-to-end operating workflow connecting the parts (read the paper → search the references by capability → classify concerns → structure feedback), and there are no validation checkpoints (e.g., verifying a suspected confounder against the paper's methods before flagging it as critical). Not a destructive/batch skill, so no cap applies; anchor 3 (steps present but checkpoints implicit) is the best fit.

3 / 5

Progressive Disclosure

The body is a genuine overview with all detail pushed to seven real, one-level-deep reference files that all exist and are clearly linked with per-file descriptions — close to anchor 5. It drops to anchor 4 because the reference files are listed in two separate sections ('Core Capabilities' and 'Reference Materials') with overlapping descriptions, a minor organization gap that creates two competing navigation lists.

4 / 5

Total

14

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, third-person description that clearly states what the skill does, when to use it, and which frameworks it applies, with explicit boundary guidance toward the peer-review skill. The only weaknesses are a mildly redundant 'Best for...' clause and a few missing natural synonyms.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'Evaluate scientific claims and evidence quality', 'assessing experimental design validity', 'identifying biases and confounders', 'applying evidence grading frameworks (GRADE, Cochrane Risk of Bias)', 'teaching critical analysis' — which is comprehensive coverage of the skill's domain. It exceeds anchor 4 (several specific actions, minor gaps) because every capability area of the body is named explicitly with the specific frameworks used; the only soft spot is the mildly redundant phrase 'Best for understanding evidence quality, identifying flaws', which repeats rather than adds.

5 / 5

Completeness

Both 'what' ('Evaluate scientific claims and evidence quality... applying evidence grading frameworks') and 'when' ('Use for assessing experimental design validity... or teaching critical analysis') are explicitly stated with concrete trigger phrases, matching the anchor-5 example pattern. A cap at 3 would only apply if the 'Use for...' clause were absent; here it is explicit.

5 / 5

Trigger Term Quality

Natural trigger phrases are well covered: 'experimental design validity', 'biases and confounders', 'GRADE', 'Cochrane Risk of Bias', 'evidence quality', 'critical analysis' — terms a user assessing a study would plausibly say. It falls short of anchor 5 because common variations like 'research methodology', 'statistical validity', 'study critique', or 'risk of bias' (the exact acronym users often type, ROB) are missing.

4 / 5

Distinctiveness Conflict Risk

The niche is clear (critical evaluation of scientific claims/evidence) and it even disambiguates the nearest neighbor: 'For formal peer review writing use peer-review.' Anchor 5 (clear niche with distinct triggers, minimal conflict risk) fits; overlap risk with generic analysis skills is low because the triggers are domain-specific.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
K-Dense-AI/claude-scientific-writer
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.