CtrlK
BlogDocsLog inGet started
Tessl Logo

scholar-evaluation

Implements the ScholarEval framework to evaluate scholarly documents; trigger when the user provides a PDF/DOCX/TXT file or pasted text and requests critique, scoring, or quality assessment.

58

Quality

73%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./scientific-skills/Evidence Insight/scholar-evaluation/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

56%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-organized with concrete commands, a clear two-path workflow (file vs. pasted text), and proper use of one-level-deep references. Its real defects are executable-correctness and redundancy: the documented scores.json key format does not match the bundled calculator's expected keys, a referenced requirements.txt is missing from the bundle, and the dimension list/commands are repeated across sections that duplicate the reference file.

Suggestions

Fix the scores.json example to use the exact keys scripts/calculate_scores.py expects ("Problem Formulation", "Methodology", ...) — or normalize keys in the script — so the documented example works verbatim.

Remove the broken requirements.txt pointer (the file is absent) and instead inline the pinned dependencies or add the file to the bundle.

State the 8-dimension list and 1–5 scale once, deferring both to references/evaluation_framework.md, and show the extraction command once; add a checkpoint to verify extraction output and check the calculator's warnings for missing dimensions.

DimensionReasoningScore

Conciseness

The body is mostly lean and does not explain known concepts, but includes noticeable redundancy that could be tightened: the 8-dimension list appears in 'Key Features', again in full under 'Implementation Details', and again in references/evaluation_framework.md; the extraction command is shown three times; and the 1–5 scale plus 'The extraction script is designed to locate the file even if the full path is not provided' restate content that belongs in the reference file. This matches anchor 3 ('mostly efficient but includes some unnecessary explanation or could be tightened') rather than anchor 4, where such duplication would be only minor.

3 / 5

Actionability

Commands are copy-paste ready (extract_text.py, calculate_scores.py) and the scores JSON example covers all 8 dimensions, but the flagship example is factually broken against the bundle: SKILL.md's scores.json uses snake_case keys ("problem_formulation") while scripts/calculate_scores.py matches on "Problem Formulation"-style keys, so following the example verbatim yields total_weight 0 and a final score of 0.0 with warnings. The body also points to a requirements.txt that is absent from the bundle. This lands on anchor 3 ('missing key details') rather than anchor 4 ('minor gaps'), because the documented example fails silently when executed as written.

3 / 5

Workflow Clarity

Example A gives a clear numbered sequence (extract → write scores.json → calculate → produce report) with concrete commands at each step, and Example B correctly short-circuits extraction for pasted text. It stops short of anchor 5 because there are no validation checkpoints: nothing tells the user to confirm extraction produced usable text, to check the script's stderr warnings for missing dimensions, or what to do when the computed score is 0/unexpected — the report step is a one-line instruction with no feedback loop.

4 / 5

Progressive Disclosure

Structure is reasonable: SKILL.md acts as an overview and points to references/evaluation_framework.md (one level deep, signaled in two places) and the two scripts, which all exist. But scored against the actual bundle, the referenced `requirements.txt` does not exist (a broken pointer in 'Dependencies'), and the 8-dimension list and 1–5 scale are inlined in SKILL.md even though they already live in the reference file — content that should be separate is inline. That combination matches anchor 3 rather than anchor 4's 'minor organization gaps'.

3 / 5

Total

13

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, well-formed description: it states a clear what, an explicit when with concrete trigger phrases (file types, pasted text, critique/scoring requests), and stays concise in third person. Its weaknesses are mild — the capability verbs are generic, and a few natural synonyms (manuscript, paper, peer review) would improve trigger coverage.

Suggestions

Enumerate one or two concrete capabilities in the description (e.g., 'extracts text and produces a weighted 8-dimension score report') to lift specificity from generic verbs to concrete actions.

Add common user synonyms such as 'paper', 'manuscript', 'thesis', or 'peer review' to the trigger phrase list for broader natural-keyword coverage.

DimensionReasoningScore

Specificity

The description names the domain ("evaluate scholarly documents") and a small set of actions ("critique, scoring, or quality assessment"), matching the anchor 'Names domain and 1-2 concrete actions, but not comprehensive'. It stays at the level of generic verbs (evaluate/critique/score) without concrete capabilities such as text extraction, weighted scoring, or report generation, so it does not reach anchor 4's 'several specific actions'.

3 / 5

Completeness

It explicitly answers both questions: what ("Implements the ScholarEval framework to evaluate scholarly documents") and when ("trigger when the user provides a PDF/DOCX/TXT file or pasted text and requests critique, scoring, or quality assessment") with concrete trigger phrases. This matches anchor 5 exactly; anchor 4 would require the 'when' to be less explicit than it is here.

5 / 5

Trigger Term Quality

It covers natural trigger terms users would say — "PDF/DOCX/TXT file", "pasted text", "critique, scoring, or quality assessment" — giving good keyword coverage with file extensions included. A few natural synonyms are missing (e.g., "review my paper", "manuscript", "peer review", "thesis"), which keeps it below anchor 5's comprehensive coverage including synonyms.

4 / 5

Distinctiveness Conflict Risk

The ScholarEval framing plus the critique/scoring/quality-assessment combination carves a mostly distinct niche with specific triggers. There is minor overlap risk with generic document-review or paper-feedback skills for requests like 'assess this file's quality', so it sits at anchor 4 rather than anchor 5's minimal-conflict clear niche.

4 / 5

Total

16

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
aipoch/medical-research-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.