CtrlK
BlogDocsLog inGet started
Tessl Logo

scholar-evaluation

Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls. Never use for ranking people or consequential decisions.

61

Quality

71%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/scholar-evaluation/SKILL.md

The canonical home for this skill is scholar-evaluation in K-Dense-AI/scientific-agent-skills

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The skill body is a well-structured operational overview with fully executable local tooling commands, a clearly sequenced 8-step workflow with validation and fail-closed checkpoints, and exemplary progressive disclosure into a complete, verified one-level-deep bundle. The only real cost is token weight from inline dates and a lengthy citation procedure.

DimensionReasoningScore

Conciseness

The body is mostly efficient — terse imperative lists, no explanation of concepts Claude already knows — but it carries inline time-sensitive details ("revised 2026-02-28", "verification dated 2026-07-23") outside any old-patterns section, and a long citation procedure that could be tightened. This fits anchor 3 (mostly efficient, some tightening possible) rather than 4, where the dated material would need to be trimmed or relocated.

3 / 5

Actionability

Guidance is fully executable: six copy-paste-ready `PYTHONDONTWRITEBYTECODE=1 python3 scripts/...` commands with exact paths and flags, concrete record fields for step 1, and explicit tri-state rating rules (`rated`/`missing`/`not_applicable`). All referenced template and script files exist in the bundle, so the commands run as written.

5 / 5

Workflow Clarity

The 8-step workflow is clearly sequenced with explicit validation checkpoints at each fragile point: rubric validation (step 3), traceability and process checks with a fail-closed checklist (step 6), stop conditions for prohibited contexts (steps 1 and the safety boundary), and a mandatory human verification checklist (step 8). This matches anchor 5's explicit validation steps, error-recovery (fail-closed), and checklists.

5 / 5

Progressive Disclosure

The body is an overview that pushes detail to 5 references, 5 asset templates, and 8 scripts — every referenced path was verified to exist in the bundle and is one level deep. References are clearly signaled at point of use ("Read `references/responsible_assessment.md` before any organizational use") and indexed with one-line descriptions in the Bundled resources section, matching anchor 5.

5 / 5

Total

18

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description names a clear, distinctive niche with several concrete actions and a strong negative safety boundary, but it lacks an explicit positive 'Use when...' trigger clause and omits common natural synonyms (paper, manuscript, draft, feedback). It reads as domain-expert wording rather than user-facing trigger language.

Suggestions

Add an explicit positive trigger clause, e.g. 'Use when the user asks for developmental feedback on a paper, draft, or research idea, or wants to audit a low-stakes assessment rubric.'

Include natural user-side synonyms such as 'paper', 'manuscript', 'draft', 'feedback', and 'peer review' alongside 'scholarly works' to improve trigger-term coverage.

Enumerate the concrete deliverables (criterion-level feedback, evidence manifests, agreement summaries, weight-sensitivity checks) so the action list is more comprehensive.

DimensionReasoningScore

Specificity

The description lists several concrete actions — "developmental review of scholarly works", "audit low-stakes research-assessment rubrics", "optional local quality controls" — with only minor coverage gaps (rating, agreement, and reporting activities are unmentioned). It fits anchor 4 rather than 3 because more than 1-2 actions are named, and not anchor 5 because the action wording ('qualitative-first, evidence-traceable') is somewhat abstract and not comprehensive.

4 / 5

Completeness

The 'what' is clear (developmental review of scholarly works; rubric auditing), but there is no positive 'Use when...' clause — only the negative trigger "Never use for ranking people or consequential decisions". Per the judging guidelines, a missing explicit positive trigger caps completeness at 3; the clear 'what' rules out 2.

3 / 5

Trigger Term Quality

Relevant keywords like "review", "rubric", and "scholarly works" appear, but the natural phrases a user would actually say — "paper", "manuscript", "draft", "feedback", "peer review" — are missing. This matches anchor 3 (some relevant keywords, missing common variations) rather than 4, which would require broader natural-term coverage.

3 / 5

Distinctiveness Conflict Risk

The niche is specific (qualitative scholarly-work review and low-stakes rubric auditing) and the explicit anti-trigger ("Never use for ranking people or consequential decisions") reduces misfiring. It is not 5 because there is minor overlap risk with general manuscript-feedback or code-review skills given the absence of positive trigger phrases.

4 / 5

Total

14

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
K-Dense-AI/claude-scientific-writer
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.