CtrlK
BlogDocsLog inGet started
Tessl Logo

tooluniverse-literature-deep-research

Conduct comprehensive literature research with target disambiguation, evidence grading, and structured theme extraction. Creates a detailed report with mandatory completeness checklist, biological model synthesis, and testable hypotheses. For biological targets, resolves offic...

52

Quality

65%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./scientific-skills/Evidence Insight/tooluniverse-literature-deep-research/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

58%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers strong workflow clarity (explicit phases, fallback chains, validation checklist) and highly actionable, tool-specific guidance, but it is a ~1,050-line monolith with no progressive disclosure — report templates, bibliography formats, and tool catalogs that belong in reference files are all inlined, with notable duplication. Conciseness is the weakest dimension: significant padding and stub code inflate token cost.

Suggestions

Move the full report template, bibliography file format, and tool catalog into reference files (e.g., references/report-template.md, references/tool-catalog.md), keeping SKILL.md as an overview with clearly signaled links.

Deduplicate the tool listings (Phase 2.2 vs. 'Quick Reference: Tool Categories') into a single catalog, and trim the advantage/limitation checkmark grids to one-line notes.

Replace stub code (loops ending in `pass`, 'Attempt 1: Call tool') with concrete, complete call examples including argument values.

DimensionReasoningScore

Conciseness

The ~1,050-line body is noticeably verbose: a ~270-line report template fully inlined, tool catalogs listed twice (Phase 2.2 and the 'Quick Reference' section), Python snippets that are comment/`pass` stubs, and padded advantage/limitation checkmark lists. Matches anchor 2's 'several unnecessary explanations or padded sections' rather than 3's 'mostly efficient', since substantial material duplicates or restates what Claude can infer.

2 / 5

Actionability

Concrete tool names, exact query syntax ('"[GENE_SYMBOL]"[Title] AND (mechanism OR function OR structure)'), fallback chains, a tier decision matrix, and complete output templates give mostly executable guidance. Not 5 because several code examples are stubs ('Attempt 1: Call tool', loops ending in `pass`) rather than copy-paste-ready calls.

4 / 5

Workflow Clarity

Phases 0-3 are clearly sequenced with explicit validation and feedback loops: a retry strategy with waits and fallback tools, per-fallback tables, a tiered full-text verification strategy with a decision matrix, and a 'Completeness Checklist (Verify Before Delivery)'. This matches anchor 5's 'explicit validation steps; feedback loops for error recovery; checklists for complex processes'.

5 / 5

Progressive Disclosure

A monolithic single file with no bundle files at all; the 270-line report template, bibliography JSON format, and tool catalog clearly belong in separate reference files (anchor 2's inlined-content criterion). It is not 3 because the problem is not a single 200-line inline block plus a buried reference — the entire payload is inline with zero reference files for a skill of this size.

2 / 5

Total

13

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description clearly communicates a multi-part biomedical literature research capability with concrete deliverables, but it is truncated mid-sentence and contains no 'Use when...' trigger guidance, capping completeness at 3. Its natural-language trigger coverage ('literature research', 'biological targets') is adequate but missing common variations users would say.

Suggestions

Complete the truncated sentence and append an explicit 'Use when...' clause (e.g., 'Use when the user asks for a literature review, deep research on a gene/protein target, or verification of a biomedical factoid').

Add natural trigger phrases users would actually say — 'literature review', 'find papers on', 'PubMed search', 'deep dive on a gene' — alongside the technical terms.

Clarify when to prefer this over a generic research skill (e.g., only for biological targets or biomedical questions) to sharpen distinctiveness.

DimensionReasoningScore

Specificity

Names several concrete actions ('target disambiguation, evidence grading, and structured theme extraction', 'mandatory completeness checklist, biological model synthesis, and testable hypotheses'), matching the 'several specific actions; minor gaps' anchor. Not 5 because the description is truncated mid-sentence ('resolves offic...') and 'comprehensive literature research' is generic phrasing leaving coverage gaps.

4 / 5

Completeness

The 'what' is clearly stated (conducts literature research, creates a detailed report), but there is no 'Use when...' clause and no explicit trigger guidance, capping completeness at 3 per the rubric. The truncated 'For biological targets, resolves offic...' only weakly implies when to use it, so it fits anchor 3 rather than 4.

3 / 5

Trigger Term Quality

'literature research' and 'biological targets' are natural user terms, but 'target disambiguation', 'evidence grading', and 'structured theme extraction' are jargon users would not say, and common variations like 'literature review', 'papers', or 'PubMed' are absent. Matches the 'some relevant keywords but missing common variations' anchor rather than 4's 'good keyword coverage'.

3 / 5

Distinctiveness Conflict Risk

Biomedical-specific deliverables ('biological model synthesis', 'testable hypotheses', 'evidence grading') give it a distinct niche with minor overlap risk against generic literature-search or web-research skills. Not 5 because 'comprehensive literature research' and 'report' are broad enough to collide with general research skills.

4 / 5

Total

14

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (1064 lines); consider splitting into references/ and linking

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
aipoch/medical-research-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.