CtrlK
BlogDocsLog inGet started
Tessl Logo

scholar-evaluation

Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls. Never use for ranking people or consequential decisions.

64

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/scholar-evaluation/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

85%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is highly actionable and well-structured with clear validation checkpoints and a clean one-level-deep reference layout; the main weakness is moderate verbosity in policy and interpretation sections that could be condensed.

Suggestions

Tighten the 'Metric and prestige policy' and 'Interpretation rules' sections into terser bullet lists to reduce token overhead while preserving the cautions.

Consider moving the detailed 'Data boundary' classifications and script constraints into references/local_tooling.md, keeping only the essential do-nots inline.

DimensionReasoningScore

Conciseness

The body is dense and avoids generic concepts Claude already knows, but the metric/prestige policy and interpretation-rules sections are verbose and could be tightened, so it sits at the mostly-efficient-but-could-be-tighter anchor rather than fully lean.

2 / 3

Actionability

Every workflow step includes concrete, copy-paste-ready bash commands (validate_rubric.py, calculate_scores.py, check_traceability.py, etc.) with exact arguments, matching the fully-executable anchor.

3 / 3

Workflow Clarity

The eight-step workflow is explicitly sequenced with validation checkpoints in step 6, a fail-closed process checklist, and a human-review checklist in step 8, satisfying the clear-sequence-with-explicit-validation anchor.

3 / 3

Progressive Disclosure

SKILL.md is an overview with a Bundled resources section giving one-line descriptions of one-level-deep references and assets, and all referenced paths in references/, assets/, and scripts/ resolve to real files, matching the well-signaled single-level anchor.

3 / 3

Total

11

/

12

Passed

Description

67%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and distinctive with concrete actions and a clear safety carve-out, but lacks an explicit positive "Use when..." trigger and leans on niche terminology over natural user phrasing.

Suggestions

Add an explicit 'Use when...' clause stating when to invoke the skill (e.g., reviewing a draft paper or auditing an assessment rubric).

Replace or supplement jargon like 'evidence-traceable developmental review' with natural phrases a user would say, such as 'review a paper' or 'audit a rubric'.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — "developmental review of scholarly works", "audit low-stakes research-assessment rubrics", "optional local quality controls" — matching the score-3 anchor for several specific actions rather than the single-domain score-2 anchor.

3 / 3

Completeness

The "what" is clear, but there is no explicit "Use when..." clause; the only trigger guidance is the negative "Never use for ranking people or consequential decisions", which per the judging guidelines caps completeness at 2.

2 / 3

Trigger Term Quality

Natural terms like "scholarly works" and "review" appear, but the bulk of the phrasing ("evidence-traceable", "low-stakes research-assessment rubrics") is niche jargon missing common user variations, so it lands at the some-relevant-keywords anchor rather than full coverage.

2 / 3

Distinctiveness Conflict Risk

The combination of "scholarly works", "developmental review", and an explicit prohibition on ranking people carves a clear niche unlikely to overlap with other skills.

3 / 3

Total

10

/

12

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

Total

15

/

16

Passed

Repository
K-Dense-AI/scientific-agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.