CtrlK
BlogDocsLog inGet started
Tessl Logo

scientific-critical-thinking

Evaluate scientific claims and evidence quality. Use for assessing experimental design validity, identifying biases and confounders, applying evidence grading frameworks (GRADE, Cochrane Risk of Bias), or teaching critical analysis. Best for understanding evidence quality, identifying flaws. For formal peer review writing use peer-review.

71

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

77%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a thorough, well-sequenced analytical framework with genuinely actionable checklists, but it overstays its token budget: an off-topic schematics section and an inline duplication of reference-file material dilute conciseness and progressive disclosure.

Suggestions

Remove or relocate the "Visual Enhancement with Scientific Schematics" section — it promotes an unrelated skill and a non-existent `scripts/generate_schematic.py`, adding padding without aiding the core critical-thinking workflow.

Trim the inline bias/fallacy/statistical taxonomy lists in SKILL.md to overviews and defer full catalogs to the existing `references/` files, since those files already cover the same material; this would lift both conciseness and progressive_disclosure toward 3.

Drop the redundant "Remember" recap section (lines 545-569), which restates principles already established in the Core Capabilities and Application Guidelines sections.

DimensionReasoningScore

Conciseness

The bulk is useful domain checklists rather than basic-concept explanation, but the "Visual Enhancement with Scientific Schematics" section is off-topic padding that references a non-existent "Nano Banana Pro"/`scripts/generate_schematic.py`, and the closing "Remember" section restates the skill — fitting the "mostly efficient but includes some unnecessary content / could be tightened" anchor.

2 / 3

Actionability

As an instruction-only skill it provides concrete, specific guidance — named biases, GRADE/Cochrane ROB, CONSORT/STROBE/PRISMA, "Conduct a priori power analysis (specify expected effect, desired power, alpha)", "Quote the problematic claim" — which per the scoring notes is not penalized for lacking executable code.

3 / 3

Workflow Clarity

Multi-step processes are clearly sequenced (claim evaluation 1-6, design process 1-7, 5-part critique structure with Critical/Important/Minor severity tiers) and the "When Uncertain" section gives conditional if/then feedback guidance; no batch or destructive operations require hard validation gates.

3 / 3

Progressive Disclosure

Six reference files are well-signaled and one-level-deep (all exist under references/), but the ~560-line body inlines detailed bias/fallacy/statistical taxonomies that the reference files already cover, matching the "content that should be separate is inline" anchor; additionally `scripts/generate_schematic.py` is referenced but absent.

2 / 3

Total

10

/

12

Passed

Description

100%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, well-scoped description that answers what the skill does and when to invoke it, lists concrete capabilities, and even disambiguates from the adjacent peer-review skill. Minor redundancy ("understanding evidence quality, identifying flaws" restates earlier phrasing) does not materially weaken it.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — "Evaluate scientific claims and evidence quality", "assessing experimental design validity", "identifying biases and confounders", "applying evidence grading frameworks (GRADE, Cochrane Risk of Bias)", "teaching critical analysis" — matching the multi-action anchor rather than the single-domain score-2 example.

3 / 3

Completeness

Explicitly states both what ("Evaluate scientific claims and evidence quality" plus enumerated actions) and when ("Use for assessing...", "Best for...") with explicit trigger guidance, satisfying the both-what-and-when anchor.

3 / 3

Trigger Term Quality

Covers natural terms users would actually say — "scientific claims", "evidence quality", "experimental design", "biases", "confounders", "GRADE", "Cochrane Risk of Bias", "critical analysis" — with good variation rather than the sparse score-2 pattern.

3 / 3

Distinctiveness Conflict Risk

Occupies a clear niche (scientific evidence appraisal) with distinct triggers and explicit de-confliction ("For formal peer review writing use peer-review"), making wrong-skill triggering unlikely.

3 / 3

Total

12

/

12

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (570 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

referenced_paths_exist

Referenced path issues: 1 missing

Warning

Total

13

/

16

Passed

Repository
foryourhealth111-pixel/Vibe-Skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.