CtrlK
BlogDocsLog inGet started
Tessl Logo

advanced-evaluation

This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise comparison, position bias, evaluation pipelines, or automated quality assessment.

54

Quality

61%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/advanced-evaluation/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

56%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The skill delivers genuinely actionable prompt templates, bias-mitigation protocols, and worked examples that a practitioner could apply directly, with clear sequencing in its core procedures. However, it is significantly overlong — re-explaining evaluation concepts Claude already knows, duplicating guidance across sections, and citing unverifiable percentage claims — and it claims internal reference documents that do not exist as files. Splitting the examples, prompt templates, and bias catalog into real bundle files and cutting the conceptual padding would substantially improve it.

Suggestions

Cut the conceptual exposition of well-known material (bias definitions, Likert scale explanations, agreement-metric descriptions) and keep only the mitigation protocols and selection guidance that add value beyond Claude's existing knowledge.

Create actual bundle files for the claimed internal references (e.g., references/bias-mitigation.md, references/metric-selection.md) and replace the plain-name 'Internal reference' list with clearly signaled links, moving the extended examples and prompt templates into them.

Remove unverifiable statistics ('40-60%', '15-25%') unless a source is cited, and deduplicate the Guidelines, Integration, and References sections, which repeat the same points and skill listings.

DimensionReasoningScore

Conciseness

The body spends large sections explaining concepts Claude already knows ("Position Bias: First-position responses receive preferential treatment", Likert scale granularity, what Cohen's κ measures), repeats the same points in multiple sections ("Always require justification before scores - Chain-of-thought prompting improves reliability by 15-25%" appears in both Core Concepts and Guidelines; the Integration and References sections list the same related skills twice), and includes unverifiable statistics ("reduce evaluation variance by 40-60%", "improves reliability by 15-25%") and time-sensitive metadata ("Last Updated: 2024-12-24"). This is noticeably verbose with several padded sections, fitting the score-2 anchor rather than score 3 because the padding is pervasive rather than occasional.

2 / 5

Actionability

The skill provides mostly executable guidance: complete copy-paste prompt templates for direct scoring and pairwise comparison, a concrete numbered position-swap protocol with confidence calibration rules ("Passes disagree: confidence = 0.5, verdict = TIE"), and a worked end-to-end example with real JSON outputs. It falls short of score 5 because several examples use placeholders instead of runnable content ("Response A: [Technical explanation with jargon]"), the rubric-generation output is explicitly "(abbreviated)", and the pipeline architecture is only an ASCII diagram with no implementation code.

4 / 5

Workflow Clarity

Multi-step processes are clearly sequenced with most checkpoints present: the Position Bias Mitigation Protocol ("1. First pass... 2. Second pass... 3. Consistency check: If passes disagree, return TIE with reduced confidence") includes an explicit validation checkpoint and error-recovery path, and the confidence calibration rules act as feedback logic. It is not score 5 because the Evaluation Pipeline Design section presents only a static architecture diagram with no sequenced build/validate steps, and human-in-the-loop feedback is described only abstractly ("Design feedback loop to improve automated evaluation").

4 / 5

Progressive Disclosure

The body has good section structure (When to Use, Core Concepts, Examples, Guidelines) but all 450 lines are inline in a single file, and the "Internal reference" section lists documents ("LLM-as-Judge Implementation Patterns", "Bias Mitigation Techniques", "Metric Selection Guide") that are not real files — no references/, scripts/, or assets/ directories exist and the entries have no paths or links. This matches the score-3 anchor (structure present, references not clearly signaled, content that should be separate is inline); it is not score 2 because the in-file organization is genuinely strong, and not score 4 because the only internal references are unusable.

3 / 5

Total

13

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is a strong trigger-focused description with natural, specific keywords that would fire reliably when users ask about LLM-as-judge setups, output comparison, or evaluation bias. Its main weakness is that it is entirely 'when' with no explicit 'what' — the skill's actual content (protocols, prompt templates, rubric-generation patterns) is never stated — plus modest overlap risk with the foundational evaluation skill in the same collection.

Suggestions

Open with an explicit capability statement (e.g., 'Provides prompt templates, bias-mitigation protocols, and rubric-generation patterns for evaluating LLM outputs with LLM judges') before the trigger clause, so both 'what' and 'when' are explicit.

Add common user synonyms such as 'evals', 'LLM evaluation', or 'judge model' to broaden natural trigger coverage.

Sharpen differentiation from the foundational 'evaluation' skill (e.g., 'advanced/production-grade techniques' vs. foundational concepts) to reduce conflict risk within the collection.

DimensionReasoningScore

Specificity

The description lists several specific concrete activities — "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias" — with minor gaps in coverage rather than generic filler. It does not reach score 5 because it never states what the skill itself provides (e.g., patterns, protocols, prompt templates), only what a user might ask for; it stays above score 3 because far more than 1-2 concrete actions are named.

4 / 5

Completeness

The 'when' is explicit and rich ("should be used when the user asks to... or mentions..."), but the 'what' is only implied through the quoted user requests — the description never states what the skill does or contains. This mirrors the score-3 anchor (one half explicit, the other weakly implied); it is not score 4 because that anchor requires both 'what' and 'when' explicitly, and not score 2 because the implied scope is considerably more informative than the bare "Use when working with documents" example.

3 / 5

Trigger Term Quality

Good keyword coverage of natural phrases users would say: "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", plus technical terms "direct scoring", "pairwise comparison", "position bias", "evaluation pipelines". A few natural variations are missing ("evals", "LLM evaluation", "benchmark", "judge model"), which keeps it at score 4 rather than the comprehensive synonym coverage of score 5, but it is well above the sparse keyword sets of score 3.

4 / 5

Distinctiveness Conflict Risk

Terms like "LLM-as-judge", "pairwise comparison", and "position bias" carve out a fairly distinct niche with clear triggers. Minor overlap risk remains with closely related skills — the body itself says this skill "extends the foundational evaluation concepts" of an existing 'evaluation' skill and shares integration points with 'context-fundamentals' and 'tool-design' — keeping it below score 5.

4 / 5

Total

15

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
sickn33/agentic-awesome-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.