CtrlK
BlogDocsLog inGet started
Tessl Logo

scientific-critical-thinking

Evaluate scientific claims and evidence quality. Use for assessing experimental design validity, identifying biases and confounders, applying evidence grading frameworks (GRADE, Cochrane Risk of Bias), or teaching critical analysis. Best for understanding evidence quality, identifying flaws. For formal peer review writing use peer-review.

67

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured analytical skill with strong progressive disclosure (real, one-level-deep references with clear navigation) and a concrete critique workflow template. Its main weakness is conciseness: repeated reference listings, a redundant 'When to Use' section, a reiterative 'Remember' section, and some generic critique advice Claude already knows inflate the body without adding proportional value.

Suggestions

Remove the duplicate reference listing: describe each reference file once (in Reference Materials) and have Core Capabilities link to it, rather than enumerating all seven files in both sections.

Cut or shrink the 'When to Use This Skill' section and the 'Remember' section, which restate the frontmatter triggers and the Overview/Application Guidelines already covered.

Tighten the 'Be Constructive / Be Specific / Be Proportionate / Apply Consistent Standards / Consider Context' guidance to the few non-obvious points; Claude already knows to distinguish fatal vs minor flaws and to acknowledge practical constraints.

DimensionReasoningScore

Conciseness

The body is organized but carries real redundancy: the reference files are listed twice (Core Capabilities lines 67-72 and Reference Materials lines 138-148), 'When to Use This Skill' repeats the frontmatter triggers, the 'Remember' section reiterates principles already covered, and the 'Be Constructive/Specific/Proportionate' guidance is generic critique advice Claude largely already knows.

3 / 5

Actionability

Provides a concrete, copy-usable critique template (Summary/Strengths/Concerns-by-severity/Specific Recommendations/Overall Assessment), specific uncertainty phrasings ('If X was done, then Y follows; if not, then Z is concern'), and real commands (`grep -r "pattern" references/`, the schematic-generation bash). The core evaluation methodology is deferred to references, leaving minor gaps.

4 / 5

Workflow Clarity

The critique-production sequence is clearly laid out (Summary → Strengths → Concerns by severity → Specific Recommendations → Overall Assessment) with an uncertainty-handling feedback step ('ask clarifying questions', 'provide conditional assessments'). This is an analytical rather than destructive/batch skill, so the strict validation-loop cap does not apply; no explicit validate-then-proceed checkpoint keeps it below 5.

4 / 5

Progressive Disclosure

SKILL.md is an overview with well-signaled, one-level-deep markdown links to seven real reference files (verified present in references/), content is appropriately split (procedural guidance inline, detailed frameworks in references), and navigation guidance is explicit ('Load references into context when detailed frameworks are needed', grep search tip).

5 / 5

Total

16

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, third-person description that states concrete capabilities, names specific frameworks, provides explicit trigger guidance, and disambiguates from a sibling skill. Minor redundancy ('Best for understanding evidence quality, identifying flaws' reiterates earlier points) and a few missing synonyms keep it just short of perfect on trigger coverage.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'Evaluate scientific claims and evidence quality', 'assessing experimental design validity, identifying biases and confounders', 'applying evidence grading frameworks (GRADE, Cochrane Risk of Bias)', 'teaching critical analysis' — with named frameworks, giving comprehensive coverage rather than just a domain label.

5 / 5

Completeness

Clearly answers 'what' (evaluate claims, apply grading frameworks, teach critical analysis) and 'when' with explicit concrete triggers ('Use for assessing experimental design validity, identifying biases and confounders, applying evidence grading frameworks... or teaching critical analysis'), plus a disambiguation cue ('For formal peer review writing use peer-review').

5 / 5

Trigger Term Quality

Strong natural keywords ('scientific claims', 'evidence quality', 'biases and confounders', 'experimental design') plus framework names (GRADE, Cochrane), but a few common synonyms users might say ('study quality', 'research methodology', 'systematic review') are absent.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (scientific evidence-quality assessment with named GRADE/Cochrane frameworks) and explicitly disambiguates from the related peer-review skill ('For formal peer review writing use peer-review'), minimizing conflict risk.

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 1 missing

Warning

Total

15

/

16

Passed

Repository
K-Dense-AI/scientific-agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.