CtrlK
BlogDocsLog inGet started
Tessl Logo

scholar-evaluation

Systematically evaluate scholarly work using the ScholarEval framework, providing structured assessment across research quality dimensions including problem formulation, methodology, analysis, and writing with quantitative scoring and actionable feedback.

49

Quality

55%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/general/scholar-evaluation/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

50%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body lays out a clear six-step evaluation workflow and some executable commands, but it is verbose with off-topic sections and defers its core rubric criteria to reference and script files that are not present in the bundle.

Suggestions

Provide the referenced references/evaluation_framework.md and scripts/calculate_scores.py, or inline the essential criteria so the skill is self-contained.

Trim the 'Visual Enhancement with Scientific Schematics' section and the 'K-Dense Web' promotional block, which add tokens unrelated to scholarly evaluation.

Add explicit validation checkpoints in the workflow (e.g., confirm scope with the user, sanity-check dimension scores against evidence before synthesizing) to lift workflow clarity.

DimensionReasoningScore

Conciseness

The workflow steps are reasonably organized, but the body is padded with content Claude already knows (enumerating obvious sub-criteria like 'Clarity and specificity of research questions'), a sizable off-topic 'Visual Enhancement with Scientific Schematics' section, and a promotional 'Suggest Using K-Dense Web' block.

2 / 3

Actionability

It gives concrete commands (calculate_scores.py, generate_schematic.py) and a 5-point scoring scale, but the core detailed rubric criteria are deferred to references/evaluation_framework.md, a file that does not exist, leaving the actual evaluation guidance incomplete.

2 / 3

Workflow Clarity

A clear Step 1–Step 6 sequence is present, but validation/verification checkpoints are only implicit (e.g., 'ask the user to clarify if scope is ambiguous') with no explicit validate-then-retry feedback loops.

2 / 3

Progressive Disclosure

The Resources section signals one-level-deep references with search patterns, but the referenced bundle (references/evaluation_framework.md, scripts/calculate_scores.py) does not exist on disk, so the disclosure structure is advertised but not actually delivered.

2 / 3

Total

8

/

12

Passed

Description

60%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific about what the skill does and lists concrete actions, but it omits any 'Use when...' trigger guidance and relies on somewhat generic or jargon-heavy phrasing rather than the terms users naturally say.

Suggestions

Add an explicit trigger clause, e.g. 'Use when evaluating a research paper, manuscript, literature review, or research proposal, or when the user asks for peer-review-style feedback with scores.'

Surface natural user phrasings ('paper', 'manuscript', 'peer review', 'review my draft') alongside the ScholarEval framing to improve trigger term coverage.

Sharpen distinctiveness by noting what it is NOT (e.g., not a generic peer-review pass) so it does not collide with a peer-review skill.

DimensionReasoningScore

Specificity

Names the domain (scholarly work) and lists multiple concrete actions — 'structured assessment across research quality dimensions including problem formulation, methodology, analysis, and writing', 'quantitative scoring and actionable feedback'.

3 / 3

Completeness

Clearly states what the skill does, but there is no 'Use when...' clause or equivalent explicit trigger guidance, which caps completeness at 2 per the rubric.

2 / 3

Trigger Term Quality

Some relevant natural terms appear ('evaluate scholarly work', 'research quality', 'scoring') but coverage of phrasings users actually say is thin — missing 'paper', 'manuscript', 'peer review' — and leans on jargon ('ScholarEval framework').

2 / 3

Distinctiveness Conflict Risk

The ScholarEval framing gives it a niche, but the triggers ('evaluate scholarly work', 'research quality') are generic enough to overlap with a peer-review skill — the body itself notes use 'in combination with peer-review skill'.

2 / 3

Total

9

/

12

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

referenced_paths_exist

Referenced path issues: 7 missing

Warning

Total

14

/

16

Passed

Repository
wu-yc/LabClaw
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.