CtrlK
BlogDocsLog inGet started
Tessl Logo

scholar-evaluation

Systematically evaluate scholarly and research work using the ScholarEval framework. Use when assessing academic papers, research proposals, literature reviews, or scholarly writing for quality, rigor, and publication readiness. Triggers: evaluate paper, scholar evaluation, research quality assessment, peer review scoring, publication readiness, academic paper review, rate research quality, ScholarEval.

59

Quality

70%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/documentation/research/scholar-evaluation/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

48%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The skill has a genuinely sound evaluation methodology — clear sequencing, stage-appropriate mindset principles, and excellent evidence-based anti-pattern examples — but it is padded with off-topic and duplicative material, and its executable guidance is undermined by references to scripts that are not in the bundle. Tightening the body to point at the existing reference instead of restating it would materially improve it.

Suggestions

Remove the 'Visual Enhancement with Scientific Schematics' section (~30 lines): it is off-topic for evaluation and depends on a script and external skill that are not in the bundle.

Either implement 'scripts/calculate_scores.py' or delete its two usage blocks — the body gives executable commands for a script the References section itself flags as 'not yet implemented'.

Collapse the eight dimensions' sub-bullets and the duplicated search-pattern list into 'references/evaluation_framework.md', keeping only dimension names and a one-line pointer in SKILL.md; also delete the empty bash block at the top of the Evaluation Workflow section.

DimensionReasoningScore

Conciseness

Noticeably verbose across several sections: the ~30-line 'Visual Enhancement with Scientific Schematics' digression, an empty bash block holding only comments, 8 dimensions x 4-5 generic sub-bullets restating what Claude already knows, and platitudinous 'Best Practices' items. Not level 1 because the anti-patterns and workflow sections do carry real, non-padded content.

2 / 5

Actionability

Concrete artifacts exist (5-point scale definitions, evidence-based GOOD/BAD scoring examples, a real reference file with search patterns), but both script commands ('scripts/calculate_scores.py', 'scripts/generate_schematic.py') point to files absent from the bundle — the References section itself admits '(not yet implemented)' — and much guidance stays descriptive. Falls between anchors 3 and 4; the broken executables keep it at 3.

3 / 5

Workflow Clarity

The six-step evaluation workflow (scope definition, dimension evaluation, scoring, synthesis, feedback, contextual adjustment) is clearly sequenced with a clarify-with-user checkpoint in Step 1. Not 5 because there are no explicit validate/fix/retry loops on the evaluation output; not 3 because the sequence and per-step deliverables are explicit rather than implicit.

4 / 5

Progressive Disclosure

Scored against the actual bundle: 'references/evaluation_framework.md' is real, one level deep, and clearly signaled, but 'scripts/calculate_scores.py' and the external 'scientific-schematics' skill are referenced repeatedly yet do not exist, and the inlined 8-dimension sub-bullets plus search patterns duplicate the reference file's content. Falls between anchors 3 and 4; the unresolved paths and duplication place it at 3.

3 / 5

Total

12

/

20

Passed

Description

91%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete domain coverage, an explicit use-when clause, and a rich natural-language trigger list. The only weaknesses are mildly generic action verbs and minor overlap potential with peer-review-adjacent skills.

DimensionReasoningScore

Specificity

Names the domain and several concrete facets — 'quality, rigor, and publication readiness' across 'academic papers, research proposals, literature reviews, or scholarly writing' — but the verbs remain generic ('evaluate', 'assessing') rather than a comprehensive list of distinct actions.

4 / 5

Completeness

Explicitly answers both 'what' ('Systematically evaluate scholarly and research work using the ScholarEval framework') and 'when' ('Use when assessing academic papers...'), augmented with concrete trigger phrases.

5 / 5

Trigger Term Quality

The explicit trigger list ('evaluate paper, scholar evaluation, research quality assessment, peer review scoring, publication readiness, academic paper review, rate research quality, ScholarEval') provides comprehensive natural-phrase coverage including synonyms a user would actually say.

5 / 5

Distinctiveness Conflict Risk

ScholarEval is a clear niche with distinct triggers, but 'peer review scoring' and 'publication readiness' could overlap with a general peer-review or writing skill; minor conflict risk only.

4 / 5

Total

18

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

referenced_paths_exist

Referenced path issues: 5 missing

Warning

Total

14

/

16

Passed

Repository
pantheon-org/tekhne
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.