CtrlK
BlogDocsLog inGet started
Tessl Logo

scientific-critical-thinking

Evaluate scientific claims and evidence quality. Use for assessing experimental design validity, identifying biases and confounders, applying evidence grading frameworks (GRADE, Cochrane Risk of Bias), or teaching critical analysis. Best for understanding evidence quality, identifying flaws. For formal peer review writing use peer-review.

73

Quality

92%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

85%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-structured, actionable instruction skill with clear sequencing, severity-tiered checklists, and exemplary progressive disclosure to six real reference files. Its main weakness is conciseness: the Overview, 'When to Use', 'Apply when', and 'Remember' sections overlap and restate content already present in the capability sections.

Suggestions

Remove overlap between the Overview, 'When to Use This Skill', and the per-section 'Apply when' lists — keep triggers in one place and let the capability sections define scope.

Trim the 'Remember' section, which restates principles already covered in 'Application Guidelines' and the capability bodies; consolidate or cut to reduce duplication.

Consider collapsing the 'When Uncertain' and 'When Providing Critique' sub-guidance into the relevant capability sections to avoid repeating general critique advice across multiple top-level sections.

DimensionReasoningScore

Conciseness

The Overview restates the description and the 'When to Use', 'Apply when', and 'Remember' sections substantially duplicate the capability bodies, so the body could be tightened despite assuming Claude's prior knowledge.

2 / 3

Actionability

Provides concrete directive checklists and names exact tools and standards (Cochrane RoB 2, ROBINS-I, Newcastle-Ottawa, Bonferroni/FDR, CONSORT/STROBE/PRISMA, GRADE upgrade/downgrade factors), giving actionable guidance for an instruction-only skill.

3 / 3

Workflow Clarity

Each capability is a clearly numbered sequence and 'When Providing Critique' specifies a 5-step output structure with severity-tiered (Critical/Important/Minor) checklists, plus conditional guidance in 'When Uncertain'.

3 / 3

Progressive Disclosure

The body is an overview pointing to six verified one-level-deep reference files, each signaled inline at the end of its section and consolidated in a dedicated 'Reference Materials' section with navigation guidance.

3 / 3

Total

11

/

12

Passed

Description

100%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific, complete, and distinctive: it enumerates concrete actions with explicit triggers, names assessment frameworks, and proactively disambiguates from the peer-review skill via a redirect. It hits the top anchor on every dimension.

DimensionReasoningScore

Specificity

Lists multiple concrete actions (assess experimental design validity, identify biases and confounders, apply GRADE/Cochrane ROB, teach critical analysis) and names specific frameworks, matching the 'lists multiple specific concrete actions' anchor.

3 / 3

Completeness

Explicitly states what the skill does and includes an explicit 'Use for...' trigger clause plus 'Best for...' guidance, satisfying both the 'what' and 'when' requirements.

3 / 3

Trigger Term Quality

Covers both lay terms ('scientific claims', 'evidence quality') and framework-specific triggers ('GRADE', 'Cochrane Risk of Bias', 'confounders', 'critical analysis') that users would naturally say.

3 / 3

Distinctiveness Conflict Risk

Occupies a clear scientific-evidence-quality niche and explicitly redirects formal peer-review writing to the separate peer-review skill, preventing overlap.

3 / 3

Total

12

/

12

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (560 lines); consider splitting into references/ and linking

Warning

referenced_paths_exist

Referenced path issues: 1 missing

Warning

Total

14

/

16

Passed

Repository
K-Dense-AI/scientific-agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.