CtrlK
BlogDocsLog inGet started
Tessl Logo

scholar-evaluation

Systematically evaluate scholarly and research work using the ScholarEval framework. Use when assessing academic papers, research proposals, literature reviews, or scholarly writing for quality, rigor, and publication readiness. Triggers: evaluate paper, scholar evaluation, research quality assessment, peer review scoring, publication readiness, academic paper review, rate research quality, ScholarEval.

55

Quality

63%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./crates/skill-auditor/tests/golden-corpus/fixtures/grade-cplus-scholar-evaluation/SKILL.md

The canonical home for this skill is scholar-evaluation in pantheon-org/tekhne

SKILL.md
Quality
Evals
Security

Quality

Content

48%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body has a genuinely clear six-step methodology and strong evidence-based anti-patterns, but it is padded with repetitive and off-topic sections, inlines detail that belongs in the existing reference file, and directs the user to scripts that are missing or unimplemented. Tightening the body to the workflow plus anti-patterns and deferring dimension detail to the reference would lift every dimension.

Suggestions

Cut the 'Visual Enhancement with Scientific Schematics' section, the comment-only bash block, and the 'Best Practices' list (it duplicates the Mindset principles), cutting the body roughly in half.

Replace the inlined 8-dimension checklists with a one-line-per-dimension summary pointing to `references/evaluation_framework.md`, which already contains the detailed rubrics.

Remove or actually implement the script references — the `python scripts/calculate_scores.py --scores ...` command cannot run against a '(not yet implemented)' script, and `scripts/generate_schematic.py` does not exist in the bundle.

DimensionReasoningScore

Conciseness

The ~350-line body has several padded sections: an empty bash block containing only comments, a 'When to Use' list of 8 bullets that restates the frontmatter triggers, a 'Best Practices' list of 8 bullets that largely repeats the three Mindset principles, and a lengthy off-topic 'Visual Enhancement with Scientific Schematics' section. This matches 'noticeably verbose; several unnecessary explanations or padded sections'; it is not score 1 because it never explains concepts Claude already knows (no 'what is peer review' filler).

2 / 5

Actionability

There is concrete guidance — a 5-point scoring scale with defined anchors, BAD/GOOD evidence examples in Anti-Patterns, and a worked example workflow — but key executable pieces are broken: `scripts/calculate_scores.py` is given a usage command yet is '(not yet implemented)', `scripts/generate_schematic.py` does not exist in the bundle, and the first code block is an empty placeholder. This fits 'some concrete guidance but incomplete... missing key details' rather than anchor 4's 'minor gaps'.

3 / 5

Workflow Clarity

The six-step evaluation workflow (scope definition → dimension evaluation → scoring → synthesis → feedback → contextual adjustment) is clearly sequenced, with checkpoints like 'Ask the user to clarify if the scope is ambiguous' and 'ALWAYS confirm the work type and adjust thresholds before scoring'. It is not score 5 because there is no explicit validation of the evaluation output itself (e.g., verifying every score cites evidence before synthesizing), leaving a minor validation gap.

4 / 5

Progressive Disclosure

The one real bundle file (`references/evaluation_framework.md`) is well signaled in a Resources section with search patterns, but the SKILL.md body inlines the full 8-dimension checklist that duplicates the framework reference's content, and two referenced paths (`scripts/calculate_scores.py`, `scripts/generate_schematic.py`) point at files that do not exist. This matches 'some structure but could be better organized; content that should be separate is inline' rather than anchor 4, where placement and references are mostly sound.

3 / 5

Total

12

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: explicit 'Use when...' clause, concrete trigger list with synonyms, and a clearly framed niche. The main gap is that the capability side rests on a single generic verb (evaluate/assess) rather than several distinct concrete actions.

Suggestions

Enumerate 2-3 concrete sub-actions in the what-clause (e.g., 'scores work across 8 quality dimensions, generates evidence-based improvement recommendations, and benchmarks publication readiness') instead of the single generic 'evaluate/assess'.

Add one or two natural user phrasings to the trigger list, such as 'review my manuscript' or 'is this paper ready to submit', to close the remaining keyword gap.

DimensionReasoningScore

Specificity

The description names the domain ('scholarly and research work') and a concrete action ('Systematically evaluate... using the ScholarEval framework') with evaluation aspects ('quality, rigor, and publication readiness'), but the action set is really one generic verb — evaluate/assess — applied to many object types, with no concrete sub-actions like dimension scoring or structured feedback reports. It sits at the 'domain and 1-2 concrete actions, not comprehensive' anchor; score 4 would require several distinct specific actions.

3 / 5

Completeness

It explicitly answers both parts: what ('Systematically evaluate scholarly and research work using the ScholarEval framework... for quality, rigor, and publication readiness') and when ('Use when assessing academic papers, research proposals, literature reviews, or scholarly writing'), plus a concrete trigger-phrase list. This matches the anchor 'clearly and explicitly answers both what AND when with concrete trigger phrases'; it is not score 4 because the 'when' clause is already explicit and specific.

5 / 5

Trigger Term Quality

The trigger list covers good natural phrasing with synonyms — 'evaluate paper', 'peer review scoring', 'rate research quality', 'publication readiness', 'academic paper review', 'ScholarEval'. A few natural user phrasings are missing (e.g., 'review my manuscript', 'is this paper good enough to submit'), so it lands at 'good keyword coverage; a few natural terms missing' rather than the comprehensive anchor 5.

4 / 5

Distinctiveness Conflict Risk

The ScholarEval framing and trigger terms ('scholar evaluation', 'ScholarEval', 'publication readiness') carve a clear niche. There is minor overlap risk with a generic peer-review skill ('peer review scoring' could plausibly trigger a peer-review skill — the body itself references one), matching 'mostly distinct; minor overlap risk with closely related skills' rather than the minimal-conflict anchor 5.

4 / 5

Total

16

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

referenced_paths_exist

Referenced path issues: 5 missing

Warning

Total

14

/

16

Passed

Repository
pantheon-org/tekhne
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.