CtrlK
BlogDocsLog inGet started
Tessl Logo

scientific-critical-thinking

Evaluate scientific claims and evidence quality. Use for assessing experimental design validity, identifying biases and confounders, applying evidence grading frameworks (GRADE, Cochrane Risk of Bias), or teaching critical analysis. Best for understanding evidence quality, identifying flaws. For formal peer review writing use peer-review.

65

Quality

77%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/scientific-critical-thinking/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-structured with a clear critique workflow, concrete templates, and exemplary one-level-deep progressive disclosure to real reference files. Its chief weakness is conciseness — the Overview, When-to-Use, and Remember sections restate content already present elsewhere — and the lack of a worked end-to-end example.

Suggestions

Collapse the 'Overview' and 'When to Use This Skill' sections, which restate the frontmatter description, into a single concise trigger list to remove redundancy.

Trim the 'Remember' section, whose principles already appear under 'Application Guidelines' and 'Always distinguish between', keeping only net-new guidance.

Add one short worked example (e.g. evaluating a sample claim through the critique template) to move actionability from mostly-executable to fully copy-paste ready.

DimensionReasoningScore

Conciseness

Mostly efficient procedural guidance, but several sections restate the description — 'When to Use This Skill' repeats the frontmatter bullets, 'Overview' and 'Remember' restate principles already in 'Application Guidelines', and 'all research has limitations' recurs — so it could be tightened; not a 4 because the redundancy is more than minor.

3 / 5

Actionability

Provides a concrete critique template (Summary/Strengths/Concerns-by-severity/Recommendations/Overall), severity tiers, phrasing templates for uncertainty, and an executable schematic command, but offers no worked example of applying the framework to a paper; mostly actionable with minor gaps rather than fully copy-paste ready.

4 / 5

Workflow Clarity

The critique is laid out as a clear five-step sequence with implicit severity triage and a conditional 'When Uncertain' branch; no destructive/batch operation is present so the validation-cap does not apply, leaving only minor checkpoint gaps versus the level-5 anchor.

4 / 5

Progressive Disclosure

The body is an overview that points via markdown links to seven real one-level-deep reference files (core_capabilities.md, scientific_method.md, common_biases.md, etc., all present in references/), each summarized in the 'Reference Materials' section with grep-based navigation guidance, matching the level-5 anchor.

5 / 5

Total

16

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is strong: third-person voice, a clear what/use-for-when structure, named frameworks that act as distinct triggers, and explicit disambiguation from the peer-review skill. The main gap is that the 'when' is framed as tasks rather than concrete user-utterance trigger phrases, and a few natural synonyms are missing.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'assessing experimental design validity, identifying biases and confounders, applying evidence grading frameworks (GRADE, Cochrane Risk of Bias), or teaching critical analysis' — giving comprehensive coverage of what the skill does, matching the level-5 anchor; it is not a 4 because coverage is broad rather than having minor gaps.

5 / 5

Completeness

Explicitly answers 'what' ('Evaluate scientific claims and evidence quality') and 'when' via a 'Use for...' clause listing use cases, but the 'when' is task-oriented rather than user-utterance trigger phrases, so it is not the level-5 anchor with concrete 'when the user mentions X' phrasing.

4 / 5

Trigger Term Quality

Includes natural terms a user would say — 'scientific claims', 'evidence quality', 'biases and confounders', 'GRADE', 'Cochrane Risk of Bias', 'peer review' — but omits common synonyms present in the body (e.g. 'systematic review', 'meta-analysis'), so it sits at good-but-not-comprehensive coverage rather than 5.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (scientific evidence-quality evaluation, GRADE/Cochrane ROB) and explicitly disambiguates the closest neighbor — 'For formal peer review writing use peer-review' — giving minimal conflict risk, matching the level-5 anchor.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
K-Dense-AI/scientific-agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.