CtrlK
BlogDocsLog inGet started
Tessl Logo

agent-evaluation

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks Use when: agent testing, agent evaluation, benchmark agents, agent reliability, test agent.

73

1.05x
Quality

60%

Does it follow best practices?

Impact

99%

1.05x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./cli-tool/components/skills/ai-research/agent-evaluation/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

33%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a concept catalog rather than an actionable skill: it names useful patterns and anti-patterns but provides no code, commands, sequenced workflow, or validation steps, and even contains a truncated sentence and placeholder '// ...' solutions. Its only real strength is compact, clearly-headed organization.

Suggestions

Add a concrete, sequenced evaluation workflow (e.g., 1. define behavioral invariants 2. run N times 3. aggregate distribution 4. validate against thresholds) with an explicit validate/fix/retry feedback loop, since batch test runs require validation checkpoints.

Replace the one-line Pattern descriptions and '// comment' Sharp Edges solutions with executable examples (e.g., a short pytest/statistics snippet for statistical test evaluation and an invariant-check snippet for behavioral contract testing).

Remove the motivational roleplay intro and fix the truncated sentence ('the goal isn't 100% test pass rate—it') so the body opens directly with actionable guidance.

DimensionReasoningScore

Conciseness

Mostly terse (lists and one-liners), but the two opening roleplay paragraphs ('You're a quality engineer who has seen agents that aced benchmarks fail spectacularly in production...') restate motivation Claude already knows and could be trimmed; matches 'mostly efficient but includes some unnecessary explanation'.

3 / 5

Actionability

Patterns are high-level hints with no executable detail ('Run tests multiple times and analyze result distributions', 'Define and test agent behavioral invariants') and Sharp Edges 'Solution' cells are '// comment' placeholders rather than real guidance; minimal concrete guidance, missing the steps to execute, matching the score-2 anchor.

2 / 5

Workflow Clarity

There is no sequenced workflow at all — the body is a categorical catalog (Capabilities/Patterns/Anti-Patterns) with no numbered steps, and no validation/feedback loop for batch test runs; matches 'steps missing or incoherent; no sequence; no validation for risky operations'.

1 / 5

Progressive Disclosure

The body is short and self-contained under well-labeled section headers with no need for external files (none provided), giving good structure; not 5 because bare-label sections like Anti-Patterns ('❌ Single-Run Testing' with no body) leave minor organization gaps.

4 / 5

Total

10

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description clearly conveys both capability and trigger context with strong, natural trigger terms and a distinct niche. Its main weakness is an awkwardly inserted over-claim ('where even top agents achieve less than 50% on real-world benchmarks') that adds fluff without aiding discovery.

DimensionReasoningScore

Specificity

Lists several concrete actions ('behavioral testing, capability assessment, reliability metrics, and production monitoring', 'benchmarking'), matching the 'lists several specific actions; minor gaps' anchor; not 5 because of the over-claim padding 'where even top agents achieve less than 50% on real-world benchmarks'.

4 / 5

Completeness

Explicitly answers both what ('Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring') and when ('Use when: agent testing, agent evaluation, benchmark agents, agent reliability, test agent') with concrete trigger phrases, matching the top anchor.

5 / 5

Trigger Term Quality

'Use when: agent testing, agent evaluation, benchmark agents, agent reliability, test agent' gives good, natural keyword coverage including a synonym pair (agent testing / test agent); not 5 because natural variants like 'evaluating agents' or 'agent benchmarks' are missing.

4 / 5

Distinctiveness Conflict Risk

Targets a clear niche (LLM agent evaluation/benchmarking) with distinct, specific triggers unlikely to fire for unrelated skills; minor overlap with adjacent agent skills is insufficient to drop it below the top anchor.

5 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
davila7/claude-code-templates
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.