CtrlK
BlogDocsLog inGet started
Tessl Logo

agent-evaluation

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks Use when: agent testing, agent evaluation, benchmark agents, agent reliability, test agent.

46

Quality

48%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.agent/skills/agent-evaluation/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

22%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

This skill is essentially a stub—it provides a persona narrative and a list of links to sub-skills but contains no actionable content of its own. The introductory text is cut off mid-sentence, there are no concrete examples, commands, or workflows, and the 'Patterns' section is empty. Without the sub-skill files in the bundle, this SKILL.md provides almost no standalone value.

Suggestions

Add a concrete quick-start workflow showing how to set up and run a basic agent evaluation (e.g., a numbered sequence with actual commands or code for running behavioral tests).

Replace the persona narrative with actionable content: a decision tree or table showing which evaluation approach to use for different scenarios (regression testing vs. capability assessment vs. adversarial testing).

Include at least one executable code example demonstrating a core pattern, such as a statistical test evaluation with pass/fail criteria and sample output.

Complete the truncated introduction and fill in the empty 'Patterns' section with concrete anti-patterns and recommended patterns, including specific examples.

DimensionReasoningScore

Conciseness

The introductory narrative about the persona ('You're a quality engineer who has seen agents...') is unnecessary padding that Claude doesn't need. However, the overall file is short and the sub-skill references are lean.

2 / 3

Actionability

There is no concrete, executable guidance anywhere in this skill. No code examples, no commands, no specific steps—just a persona description, a list of capabilities/requirements, and links to sub-skills. The body describes rather than instructs.

1 / 3

Workflow Clarity

There is no workflow, no sequenced steps, and no validation checkpoints. The content is purely a table of contents with no process guidance for how to actually evaluate an agent.

1 / 3

Progressive Disclosure

The skill does reference sub-skills with one-level-deep links, which is good structure. However, no bundle files were provided to verify these references exist, the overview content itself is too thin to serve as a useful entry point, and the description is cut off mid-sentence suggesting incomplete content.

2 / 3

Total

6

/

12

Passed

Description

75%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description covers a well-defined niche (LLM agent testing/benchmarking) and includes both 'what' and 'when' components, which is good. However, the capability descriptions lean toward category labels rather than concrete actions, and the trigger terms are somewhat repetitive with 'agent' prefixed to every keyword. The factual claim about '<50% on real-world benchmarks' adds context but doesn't help with skill selection.

Suggestions

Replace category labels with concrete actions, e.g., 'Designs behavioral test suites for LLM agents, runs capability benchmarks, calculates reliability metrics, and sets up production monitoring dashboards.'

Diversify trigger terms to include natural user phrasings like 'evaluate my AI agent,' 'LLM evaluation,' 'agent performance testing,' 'how good is my agent,' 'agent accuracy,' or 'compare agent models.'

DimensionReasoningScore

Specificity

Names the domain (LLM agent testing/benchmarking) and lists some action areas like 'behavioral testing, capability assessment, reliability metrics, and production monitoring,' but these read more as category labels than concrete actions (e.g., no specifics like 'run benchmark suites,' 'generate reliability reports,' or 'compare agent performance').

2 / 3

Completeness

The description clearly answers both 'what' (testing and benchmarking LLM agents with behavioral testing, capability assessment, reliability metrics, production monitoring) and 'when' (explicit 'Use when' clause with trigger terms). Both components are present and explicit.

3 / 3

Trigger Term Quality

The 'Use when' clause includes relevant terms like 'agent testing,' 'agent evaluation,' 'benchmark agents,' and 'agent reliability,' but misses common natural variations users might say such as 'LLM evaluation,' 'AI agent performance,' 'agent accuracy,' 'agent scoring,' or 'evaluate my agent.' The terms are somewhat repetitive with 'agent' appearing in every trigger.

2 / 3

Distinctiveness Conflict Risk

The focus on LLM agent testing and benchmarking is a clear, specific niche that is unlikely to conflict with general testing skills, code testing skills, or generic LLM skills. The combination of 'agent' + 'testing/benchmarking' creates a distinct trigger profile.

3 / 3

Total

10

/

12

Passed

Validation

90%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation10 / 11 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

10

/

11

Passed

Repository
Dokhacgiakhoa/antigravity-ide
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.