CtrlK
BlogDocsLog inGet started
Tessl Logo

scholar-evaluation

Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls. Never use for ranking people or consequential decisions.

64

Quality

75%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/scholar-evaluation/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is exceptionally well-structured: executable commands, a validated multi-step workflow, and clean one-level-deep references that all resolve to real files. The only weakness is inline time-sensitive bibliographic detail that would be more token-efficient in a dated reference section.

DimensionReasoningScore

Conciseness

The body is dense and assumes Claude's competence without padding, but embeds time-sensitive bibliographic detail inline ('revised 2026-02-28', arXiv version, dated ledger) that the rubric says should live in a dated/deprecated section rather than the overview.

4 / 5

Actionability

Every tooling step is a copy-paste-ready bash command with exact scripts, flags, and asset paths (e.g. validate_rubric.py, calculate_scores.py with --rubric/--evaluation args), matching the 'fully executable, copy-paste ready' anchor.

5 / 5

Workflow Clarity

The 8-step workflow is explicitly sequenced with validation checkpoints and feedback loops ('Stop on a prohibited decision context', fail-closed checklist, 'Only proceed when validation passes'), satisfying the destructive/batch validation requirement rather than triggering the cap.

5 / 5

Progressive Disclosure

The body is a clear overview with one-level-deep, well-signaled references, every cited reference resolves to a real bundle file, and a 'Bundled resources' index organizes discovery, matching the top anchor.

5 / 5

Total

19

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and distinct with a clear niche and explicit safety boundary, but it lacks an explicit 'Use when...' trigger clause and omits common user-facing trigger phrases. Adding concrete positive trigger guidance would lift completeness and trigger-term quality.

Suggestions

Add an explicit positive 'Use when reviewing a scholarly work or auditing a low-stakes assessment rubric' clause so the 'when' is stated, not implied.

Include natural user-facing trigger phrases (e.g., 'paper review', 'draft feedback', 'rubric validation') and common synonyms alongside the academic terminology.

Lead with the primary positive capability before the negative boundary so the trigger is recognisable before the prohibition.

DimensionReasoningScore

Specificity

Names the domain ('scholarly works', 'research-assessment rubrics') and several concrete actions ('developmental review', 'audit', 'local quality controls'), with only minor coverage gaps, matching the 'lists several specific actions' anchor.

4 / 5

Completeness

The 'what' is clear (developmental review and rubric audit), but there is no explicit 'Use when...' trigger clause—only a negative boundary ('Never use for...')—so 'when' is at best weakly implied, capping completeness at 3 per the rubric guideline.

3 / 5

Trigger Term Quality

Terms like 'scholarly works', 'developmental review', and 'research-assessment rubrics' are relevant but lean academic and omit common user phrasings, synonyms, or file types, fitting 'some relevant keywords but missing common variations'.

3 / 5

Distinctiveness Conflict Risk

The niche (scholarly-work review and low-stakes rubric audit) plus an explicit negative boundary make it mostly distinct with only minor overlap risk against general review skills.

4 / 5

Total

14

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
K-Dense-AI/scientific-agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.