CtrlK
BlogDocsLog inGet started
Tessl Logo

advanced-evaluation

This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise comparison, position bias, evaluation pipelines, or automated quality assessment.

68

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A thorough, highly actionable skill body with concrete prompt templates, examples, and a well-sequenced pairwise workflow supported by real reference files. Its main weaknesses are redundancy between the Guidelines and Gotchas sections and a long inlined body that could push more detail into the existing references.

Suggestions

Merge or cross-reference the Guidelines and Gotchas sections so each principle (justification-before-scores, position swap) appears once, cutting the most visible redundancy.

Move the full worked examples and/or the long Guidelines/Gotchas lists into a reference file, keeping SKILL.md as a lean overview that points to the references already present.

Add an explicit validate-and-retry step to the direct-scoring workflow so it matches the pairwise workflow's checkpoint rigor.

DimensionReasoningScore

Conciseness

The body is mostly efficient and content-rich, but the marketing-style intro ("synthesizes research from academic papers, industry practices..."), the 'Key insight' line, and significant overlap between the Guidelines and Gotchas sections (justification-before-scores and position-swap each appear in both) add redundancy. Not 4 because the redundancy and padded prose are more than 'minor'; not 2 because the bulk is genuinely useful rather than padded filler.

3 / 5

Actionability

Provides copy-paste-ready prompt templates (direct scoring and pairwise), a concrete criteria-definition pattern, scale-calibration guidance, a numbered position-swap procedure, and full JSON example outputs. Fits the fully-executable anchor with specific examples covering common cases.

5 / 5

Workflow Clarity

The pairwise workflow is clearly sequenced with an explicit consistency-check validation step and a TIE feedback path, and the pipeline layers are listed. Not 5 because the direct-scoring workflow lacks an explicit validation/feedback loop and the pipeline is presented as layers rather than a validated sequence with checkpoints.

4 / 5

Progressive Disclosure

Four real reference files exist and are linked with clear 'Read when:' triggers (one level deep), plus external research links; the body is well-sectioned. Not 5 because the ~400-line body inlines substantial material (full examples, 10 Guidelines, 8 Gotchas) that could be split into references, leaving minor organization gaps.

4 / 5

Total

16

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, trigger-rich description that uses third-person voice and explicit 'when' guidance with concrete, natural phrases. Its main weakness is fusing the 'what' and 'when' into one sentence rather than stating capabilities declaratively before the trigger clause.

Suggestions

Lead with a concise declarative capability statement (e.g., 'Builds and calibrates LLM-as-judge evaluation systems') before the 'Use when...' clause so the 'what' is as explicit as the 'when'.

Differentiate more sharply from the foundational 'evaluation' skill by signaling the advanced/production-grade scope in the description itself.

DimensionReasoningScore

Specificity

Names multiple concrete actions ("implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias") plus specific techniques (direct scoring, pairwise comparison, position bias, evaluation pipelines), giving comprehensive coverage. Not below 5 because no meaningful capability is omitted; not above since 5 is the scale max.

5 / 5

Completeness

Explicit 'when' guidance is present ("This skill should be used when the user asks to...") and the 'what' is conveyed via the action verbs, but the two are fused into a single trigger sentence rather than a clean declarative capability statement followed by a 'when' clause. Stays at 4 (not 5) because the 'what' is not stated as explicitly as the anchor-5 example which separates capability and trigger.

4 / 5

Trigger Term Quality

Covers natural phrases users actually say ("compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", "position bias", "evaluation pipelines", "automated quality assessment") with good synonym coverage. Fits the comprehensive-coverage anchor rather than 4, where only 'a few natural terms' would be missing.

5 / 5

Distinctiveness Conflict Risk

Triggers are specific (LLM-as-judge, pairwise comparison, position bias) and carve a clear niche, but the skill collection includes a foundational 'evaluation' skill, creating minor overlap risk with a closely related skill. Not 5 because of that related-skill overlap; not 3 because the triggers are far from generic.

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
neil-tessl/skill-index-evals-dev
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.