CtrlK
BlogDocsLog inGet started
Tessl Logo

llm-evaluation

Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.

75

1.75x
Quality

66%

Does it follow best practices?

Impact

91%

1.75x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./tests/ext_conformance/artifacts/agents-wshobson/llm-application-dev/skills/llm-evaluation/SKILL.md

The canonical home for this skill is llm-evaluation in wshobson/agents

SKILL.md
Quality
Evals
Security

Quality

Content

57%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A code-rich, actionable catalog of LLM evaluation techniques whose implementations are near copy-paste ready. Its weaknesses are topical rather than sequential organization, glossary sections restating what Claude already knows, and a monolithic structure with no reference files despite content that would split naturally.

Suggestions

Add an explicit evaluation workflow section that sequences the existing pieces (define test cases -> select metrics -> establish baseline -> run -> check statistical significance -> gate deployment via regression detection) so the catalog becomes a process.

Cut the 'Core Evaluation Types' and 'Human Evaluation / Dimensions' glossaries, which re-explain BLEU, ROUGE, accuracy, and precision/recall that Claude already knows; keep only non-obvious domain judgment such as the Common Pitfalls list.

Move the metric implementations, LLM-as-judge patterns, and LangSmith integration into separate reference files (e.g., references/metrics.md, references/llm-as-judge.md) linked from a lean overview SKILL.md, and make the Quick Start self-contained by defining or importing the functions it calls.

DimensionReasoningScore

Conciseness

The "Core Evaluation Types" and "Human Evaluation / Dimensions" sections re-explain concepts Claude already knows ("Accuracy: Percentage correct", "BLEU: N-gram overlap", "Precision/Recall/F1: Class-specific performance"), which is unnecessary padding. Not 2 because the bulk of the body is dense, payload-bearing code rather than padded prose.

3 / 5

Actionability

Runnable implementations for BLEU, ROUGE, BERTScore, groundedness, LLM-as-judge, Cohen's kappa, A/B testing, regression detection, and LangSmith are mostly copy-paste ready. Not 5 because Quick Start calls calculate_bleu/calculate_bertscore/check_groundedness before they are defined and relies on a your_model placeholder, and the judge snippets parse raw JSON output instead of tool-use.

4 / 5

Workflow Clarity

Content is organized by topic rather than as a sequenced evaluation workflow; there is no explicit end-to-end process (e.g., define test set -> choose metrics -> establish baseline -> run -> check significance -> gate deployment) with validation checkpoints. Not 2 because each section is individually coherent and regression detection plus significance testing supply checkpoint ingredients.

3 / 5

Progressive Disclosure

No bundle files exist and none are referenced; the ~690-line body has clear section headers but inlines metric implementations, judge patterns, and framework integration that would fit separate reference files. This matches 'some structure ... content that should be separate is inline' rather than 4, since nothing is offloaded at all.

3 / 5

Total

13

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A well-formed description that explicitly covers both capability and usage triggers in third-person imperative style. It names three concrete evaluation approaches and three natural trigger phrases, with only minor gaps in synonym coverage and slight generic wording.

DimensionReasoningScore

Specificity

"Implement comprehensive evaluation strategies ... using automated metrics, human feedback, and benchmarking" names the domain plus three concrete method categories, matching the 'several specific actions; minor gaps' anchor. It falls short of 5 because the actions are mechanisms rather than concrete tasks, and 'comprehensive' is mild fluff.

4 / 5

Completeness

Both what ("Implement comprehensive evaluation strategies ... using automated metrics, human feedback, and benchmarking") and an explicit when ("Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks") are present. Not 5 because the when clause repeats 'evaluation frameworks' from the what and lacks more concrete trigger variations.

4 / 5

Trigger Term Quality

"testing LLM performance", "measuring AI application quality", and "establishing evaluation frameworks" are phrases users would naturally say, giving good keyword coverage. Not 5 because the domain's most common natural terms ('eval', 'evals', 'LLM evaluation') and synonyms are missing.

4 / 5

Distinctiveness Conflict Risk

LLM application evaluation is a clear niche with distinct triggers ("testing LLM performance", "measuring AI application quality"), so overlap risk is minor. Not 5 because generic terms like "benchmarking" could pull in non-LLM testing contexts.

4 / 5

Total

16

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (696 lines); consider splitting into references/ and linking

Warning

Total

15

/

16

Passed

Repository
Dicklesworthstone/pi_agent_rust
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.