CtrlK
BlogDocsLog inGet started
Tessl Logo

scholar-evaluation

Systematically evaluate scholarly work using the ScholarEval framework, providing structured assessment across research quality dimensions including problem formulation, methodology, analysis, and writing with quantitative scoring and actionable feedback.

54

Quality

62%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./bundled/skills/scholar-evaluation/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

57%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The skill is well-structured with genuine progressive disclosure pointing to real bundle files, and its six-step workflow is easy to follow. Its weaknesses are token-padding verbosity, mostly descriptive rather than executable evaluation guidance, and the absence of validation checkpoints in the workflow.

Suggestions

Trim the inline dimension bullet lists and the scientific-schematics section to pointers, moving detail to references to reduce token cost.

Make the evaluation guidance executable: show the exact output schema for a dimension assessment and a concrete worked scoring example with real numbers rather than an outline.

Add explicit validation checkpoints (e.g., 'confirm every applicable dimension has a score before calculating the aggregate') and a feedback loop for inconsistent or missing ratings.

DimensionReasoningScore

Conciseness

The body is mostly organized and avoids explaining basic concepts, but it is padded with verbatim dimension bullet lists and an extended scientific-schematics section that largely restates possibilities Claude already knows, so it could be tightened considerably.

2 / 3

Actionability

It gives some concrete guidance (script invocation lines, a 5-point scoring scale, a worked example outline) but the core evaluation steps are described at a checklist level rather than as fully executable instructions, and the schematic-generation bash example references a script that is not in the bundle.

2 / 3

Workflow Clarity

The six evaluation steps are clearly sequenced, but there are no validation checkpoints or feedback loops (e.g., verifying a dimension was actually assessed before scoring, or sanity-checking an aggregate score), and a destructive/batch context cap at 2 does not apply, so this sits at the 'steps listed, checkpoints implicit' anchor.

2 / 3

Progressive Disclosure

The SKILL.md body is an overview that clearly signals one-level-deep references — references/evaluation_framework.md and scripts/calculate_scores.py are both real files, described with load-when guidance and search patterns, and detailed rubrics are appropriately offloaded to the reference file rather than inlined.

3 / 3

Total

9

/

12

Passed

Description

67%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and distinctive, clearly conveying a scholarly-evaluation niche with concrete assessment dimensions. Its main weakness is the absence of an explicit "Use when..." trigger clause and reliance on framework jargon over natural user phrasing.

Suggestions

Add an explicit trigger clause such as 'Use when evaluating research papers, literature reviews, research proposals, or scholarly writing for quality, rigor, or publication readiness.'

Replace framework-internal jargon with natural trigger terms users would actually say (e.g., 'review my paper', 'peer review', 'is this paper publishable', 'assess research quality').

Tighten the description so the concrete actions lead and the framework attribution follows, keeping it concise while preserving both what and when.

DimensionReasoningScore

Specificity

The description lists several concrete actions and dimensions — "structured assessment across research quality dimensions including problem formulation, methodology, analysis, and writing" plus "quantitative scoring and actionable feedback" — matching the multiple-specific-actions anchor rather than the single-domain score-2 example.

3 / 3

Completeness

It clearly answers "what does this do" but the "when to use it" trigger is only implied through the listed use cases; there is no explicit "Use when..." clause, which per the judging guidelines caps completeness at 2.

2 / 3

Trigger Term Quality

It contains relevant domain keywords ("scholarly work", "research", "scoring", "evaluation") but is written in framework-technical phrasing ("ScholarEval framework", "research quality dimensions") rather than the natural trigger phrases a user would say, so it has some relevant keywords but misses common variations.

2 / 3

Distinctiveness Conflict Risk

The ScholarEval framing and enumerated research-quality dimensions carve out a clear niche (academic/scholarly evaluation) with distinct triggers that are unlikely to fire for unrelated skills.

3 / 3

Total

10

/

12

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

referenced_paths_exist

Referenced path issues: 1 missing

Warning

Total

14

/

16

Passed

Repository
foryourhealth111-pixel/Vibe-Skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.