CtrlK
BlogDocsLog inGet started
Tessl Logo

scholar-evaluation

Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls. Never use for ranking people or consequential decisions.

61

Quality

71%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./scientific_writer/.claude/skills/scholar-evaluation/SKILL.md

The canonical home for this skill is scholar-evaluation in K-Dense-AI/scientific-agent-skills

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an unusually rigorous, fully actionable instruction skill: exact commands, explicit enums and file paths, a numbered workflow with validation and fail-closed checkpoints at every stage, and clean one-level-deep progressive disclosure. Its only real weakness is conciseness — governance and metric policy prose repeats material that belongs in references/responsible_assessment.md.

DimensionReasoningScore

Conciseness

The body is dense and policy-focused with no explanations of concepts Claude already knows, but it could be tightened: the metric/prestige policy detail and safety prose partially duplicate content that references/responsible_assessment.md exists to hold, and the citation section runs long. This is more than the 'minor instances' of a 4, but far from the generic padding of a 2.

3 / 5

Actionability

Guidance is fully executable: copy-paste-ready bash commands with complete argument lists for all seven scripts, exact template file paths, explicit allowed classification enums, and precise rules for rated/missing/not_applicable states including null handling. The common cases are covered end-to-end.

5 / 5

Workflow Clarity

Eight numbered steps are clearly sequenced with explicit validation checkpoints throughout: rubric structural validation before use, a fail-closed process checklist ('intentionally unconfirmed and fails closed'), stop conditions ('Stop on a prohibited decision context or unnecessary private data'), and a human review gate before release. This matches the anchor with explicit validation and feedback structure.

5 / 5

Progressive Disclosure

The 'Bundled resources' section lists every reference and asset with a one-line description, references are one level deep and clearly signaled (all cited paths verified to exist), and bulk detail (exact schemas, formulas, output behavior) is correctly deferred to references/local_tooling.md rather than inlined.

5 / 5

Total

18

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and clearly delineates a niche domain with a hard negative-use boundary, but it lacks any positive trigger guidance ('Use when...'), which caps completeness and weakens trigger-term quality. Natural user phrasings for the core tasks (paper review, draft feedback, rubric audit) are absent.

Suggestions

Add a positive trigger clause such as 'Use when the user asks for developmental feedback on a paper, draft, protocol, or literature synthesis, or wants to audit a research-assessment rubric' to raise completeness above 3.

Include natural synonyms users would say — 'paper', 'draft', 'manuscript', 'peer feedback', 'rubric audit' — to improve trigger-term coverage.

Make the 'optional local quality controls' action concrete (e.g., 'runs bundled Python validators for rubric structure, score calculation, and evidence traceability') to strengthen specificity.

DimensionReasoningScore

Specificity

The description uses third person and names the domain ('scholarly works', 'low-stakes research-assessment rubrics') plus several concrete actions ('developmental review', 'audit', 'optional local quality controls'), though 'local quality controls' is somewhat generic, leaving minor gaps in coverage rather than the comprehensive action list of a 5.

4 / 5

Completeness

The 'what' is clear (developmental review plus rubric auditing), but there is no positive 'Use when...' clause — only the negative 'Never use for ranking people or consequential decisions'. Per the judging guidelines, missing explicit positive trigger guidance caps completeness at 3.

3 / 5

Trigger Term Quality

Terms like 'scholarly works', 'research-assessment rubrics', and 'developmental review' are relevant, but common natural phrasings users would actually say — 'paper', 'draft', 'manuscript', 'peer-review feedback', 'rubric audit' — are missing, so keyword coverage has notable synonym gaps.

3 / 5

Distinctiveness Conflict Risk

The scholarly-work review and assessment-rubric audit niche is fairly distinct from typical skills with minimal overlap risk, but the absence of strong positive triggers leaves some overlap with general writing/feedback skills, so it does not fully reach the clear-niche anchor of 5.

4 / 5

Total

14

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
K-Dense-AI/claude-scientific-writer
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.