CtrlK
BlogDocsLog inGet started
Tessl Logo

advanced-evaluation

This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise comparison, position bias, evaluation pipelines, or automated quality assessment.

41

Quality

41%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/advanced-evaluation/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

27%Scale 1-3

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

This skill reads more like a comprehensive survey paper or textbook chapter than an actionable skill file. It covers the LLM-as-judge topic thoroughly but violates token efficiency by explaining concepts Claude already understands, includes no executable code for pipeline implementation, and packs everything into a single monolithic file. The prompt templates and examples provide some actionable value, but the overall verbosity and lack of progressive disclosure significantly reduce its effectiveness as a skill.

Suggestions

Cut the content by 50-60%: remove the bias taxonomy explanations (Claude knows these), the metric selection table (standard knowledge), the 'When to Use' section, and the 'Core Concepts' preamble. Focus on the novel patterns—position swap protocol, confidence calibration formula, and rubric generation templates.

Split into multiple files: move the detailed examples into an EXAMPLES.md, the bias mitigation protocols into BIAS_MITIGATION.md, and the rubric generation patterns into RUBRICS.md, keeping SKILL.md as a concise overview with clear references.

Add executable Python code for the core operations: implement the position-swap pairwise comparison as a runnable function, include a concrete confidence calibration calculation, and provide a working evaluation pipeline script rather than an ASCII diagram.

Add explicit validation checkpoints to the pipeline workflow: e.g., 'Verify inter-annotator agreement > 0.7 before deploying automated evaluation' and 'Run calibration set of 20 pre-labeled examples to validate judge accuracy before batch evaluation.'

DimensionReasoningScore

Conciseness

The skill is extremely verbose at ~350+ lines. It explains concepts Claude already knows well (what position bias is, what Spearman's ρ measures, what a Likert scale is, basic evaluation taxonomy). The 'Key insight' callout, the 'When to Use' section listing 7 bullet points, and extensive bias descriptions all add tokens without adding novel, actionable information. The ASCII pipeline diagram is decorative rather than instructive.

1 / 3

Actionability

The skill provides prompt templates and JSON output examples, which is useful. However, the prompts are not truly executable—they contain placeholders and describe patterns rather than providing copy-paste-ready code for building an actual evaluation pipeline. There's no executable code (Python, etc.) for implementing the pipeline, position swapping, or confidence calibration—just conceptual descriptions and template strings.

2 / 3

Workflow Clarity

The pairwise comparison position-swap protocol is well-sequenced with clear steps. However, the overall evaluation pipeline is presented as a conceptual diagram rather than an actionable workflow with validation checkpoints. There's no explicit verification step for checking if the evaluation system is producing reliable results before deploying it, and the pipeline design section lacks concrete error recovery guidance.

2 / 3

Progressive Disclosure

The content is a monolithic wall of text with no bundle files to reference. Everything—taxonomy, bias landscape, metric selection, implementation patterns, examples, rubric generation, scaling strategies, references—is crammed into a single file. The 'Integration' and 'References' sections mention other skills and external links but the core content that should be split (e.g., bias mitigation details, metric selection guide, rubric templates) is all inline.

1 / 3

Total

6

/

12

Passed

Description

54%Scale 1-3

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description has strong trigger term coverage and occupies a clear, distinctive niche, but it fundamentally fails at explaining what the skill actually does. It reads as a pure trigger-matching clause without any capability description, making it impossible for Claude to understand the skill's concrete functionality before selecting it.

Suggestions

Add a clear 'what it does' statement before the trigger clause, e.g., 'Implements LLM-as-judge evaluation pipelines, creates scoring rubrics, sets up pairwise comparison frameworks, and applies bias mitigation techniques for automated quality assessment.'

Restructure to lead with concrete capabilities in third person voice (e.g., 'Designs and implements...') followed by the existing 'Use when...' trigger clause.

DimensionReasoningScore

Specificity

The description names the domain (LLM-as-judge evaluation) and mentions some actions like 'implement LLM-as-judge', 'compare model outputs', 'create evaluation rubrics', 'mitigate evaluation bias', but these are embedded in trigger clauses rather than stated as concrete capabilities the skill performs. There's no clear 'what it does' statement listing specific actions.

2 / 3

Completeness

The description is essentially all 'when' with no 'what'. It tells Claude when to use the skill but never explains what the skill actually does — what concrete actions it performs, what outputs it produces, or what capabilities it provides. The 'what' is entirely missing.

1 / 3

Trigger Term Quality

Excellent coverage of natural trigger terms users would say: 'implement LLM-as-judge', 'compare model outputs', 'create evaluation rubrics', 'mitigate evaluation bias', 'direct scoring', 'pairwise comparison', 'position bias', 'evaluation pipelines', 'automated quality assessment'. These are natural phrases a user working in this domain would use.

3 / 3

Distinctiveness Conflict Risk

The description targets a very specific niche — LLM-as-judge evaluation patterns — with highly distinctive trigger terms like 'position bias', 'pairwise comparison', 'LLM-as-judge', and 'evaluation rubrics'. This is unlikely to conflict with other skills.

3 / 3

Total

9

/

12

Passed

Validation

90%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 10 / 11 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

10

/

11

Passed

Repository
popey/claude-code-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.