CtrlK
BlogDocsLog inGet started
Tessl Logo

agent-evaluation

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks

33

Quality

30%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./plugins/antigravity-bundle-agent-architect/skills/agent-evaluation/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

27%Scale 1-3

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

This skill is extremely verbose, containing hundreds of lines of illustrative but non-executable TypeScript pseudocode that explains testing concepts Claude already understands. The patterns are conceptually sound but lack practical actionability—no real tool setup, no concrete commands, and heavy reliance on undefined abstractions. The monolithic structure with no progressive disclosure makes it a poor use of context window budget.

Suggestions

Reduce content by 70-80%: replace full class implementations with concise pattern descriptions, key interfaces, and 5-10 line code snippets showing the essential logic (e.g., confidence interval calculation, flakiness formula)

Make content actionable by showing real tool usage—e.g., actual Langsmith SDK calls, AgentBench setup commands, or PromptFoo configuration files instead of abstract TypeScript classes

Split into multiple files: keep SKILL.md as a concise overview with pattern summaries, and move detailed implementations to referenced files like PATTERNS.md, SHARP_EDGES.md, and ADVERSARIAL_TESTING.md

Add explicit validation checkpoints to workflows—e.g., 'After establishing baseline, verify confidence intervals are narrow enough (CI width < 0.1) before proceeding to regression testing'

DimensionReasoningScore

Conciseness

Extremely verbose at ~600+ lines. Massive code blocks explain concepts Claude already knows (statistical testing, chi-squared tests, Jaccard similarity). The interfaces and classes are illustrative pseudocode that could be condensed to patterns and key principles. Sections like 'What is a PDF' equivalent explanations of basic testing concepts waste tokens.

1 / 3

Actionability

The code examples are TypeScript-like but not truly executable—they reference undefined types (Agent, AgentOutput, AgentContext, TestCase), unimplemented helper methods (containsRudeLanguage, isRelevantToCustomerService, containsLegalAdvice, similarity), and abstract interfaces. They illustrate patterns but aren't copy-paste ready. No concrete tool commands or real framework usage (e.g., actual Langsmith or AgentBench setup) are provided.

2 / 3

Workflow Clarity

The Collaboration section has brief workflow sequences (design → create suite → implement → evaluate → iterate), but the main patterns lack explicit validation checkpoints and feedback loops. The Statistical Test Evaluation pattern runs tests and analyzes but doesn't specify what to do when concerns are identified. The regression testing has a deploy/don't-deploy recommendation but no recovery workflow.

2 / 3

Progressive Disclosure

This is a monolithic wall of text with no references to external files despite being extremely long. All patterns, sharp edges, and collaboration details are inlined. There are no bundle files, yet the content is far too long to be effective as a single SKILL.md. Content like the full class implementations for each pattern should be in separate reference files.

1 / 3

Total

6

/

12

Passed

Description

32%Scale 1-3

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description identifies a clear domain (LLM agent testing and benchmarking) and lists relevant subcategories, but it reads more like a topic summary than an actionable skill description. It lacks a 'Use when...' clause, concrete actions the skill performs, and the trailing statistical claim ('even top agents achieve less than 50%') adds no selection value. The description needs explicit trigger guidance and more specific capability statements.

Suggestions

Add an explicit 'Use when...' clause with trigger scenarios, e.g., 'Use when the user asks about evaluating LLM agents, creating evals, measuring agent reliability, or setting up agent benchmarks.'

Replace the statistical claim with concrete actions the skill performs, e.g., 'Generates behavioral test suites, designs capability evaluations, calculates reliability metrics, and sets up production monitoring dashboards for LLM agents.'

Include common user-facing synonyms and variations like 'eval', 'evaluation', 'agent performance', 'test suite', 'accuracy measurement' to improve trigger term coverage.

DimensionReasoningScore

Specificity

The description names the domain (LLM agent testing/benchmarking) and lists some action areas (behavioral testing, capability assessment, reliability metrics, production monitoring), but these are more like categories than concrete actions. It doesn't specify what the skill actually does (e.g., 'generates test suites', 'runs benchmarks', 'produces reliability reports').

2 / 3

Completeness

The description addresses 'what' at a high level but completely lacks a 'Use when...' clause or any explicit trigger guidance for when Claude should select this skill. Per the rubric, a missing 'Use when...' clause caps completeness at 2, and the 'what' is also somewhat vague, bringing this to 1.

1 / 3

Trigger Term Quality

Includes relevant terms like 'LLM agents', 'benchmarking', 'testing', 'reliability metrics', and 'production monitoring' which users might naturally use. However, it misses common variations like 'eval', 'evaluation', 'agent evaluation', 'test harness', 'accuracy', or 'performance testing' that users would likely say.

2 / 3

Distinctiveness Conflict Risk

The focus on LLM agent testing/benchmarking is a reasonably specific niche, but the broad terms like 'testing', 'monitoring', and 'capability assessment' could overlap with general software testing skills or monitoring/observability skills. The added detail about '<50% on real-world benchmarks' is a factoid rather than a disambiguating trigger.

2 / 3

Total

7

/

12

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 9 / 11 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (1136 lines); consider splitting into references/ and linking

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

9

/

11

Passed

Repository
popey/claude-code-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.