CtrlK
BlogDocsLog inGet started
Tessl Logo

anthropic-evaluations

This skill should be used when the user asks to "create evals", "evaluate an agent", "build evaluation suite", or mentions agent testing, graders, or benchmarks. Also suggest when building coding agents, conversational agents, or research agents that need quality assurance.

86

1.50x
Quality

80%

Does it follow best practices?

Impact

98%

1.50x

Average score across 3 eval scenarios

SecuritybySnyk

Failed to scan

The risk profile of this skill

SKILL.md
Quality
Evals
Security

Quality

Content

66%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Well-structured, token-efficient overview that leverages progressive disclosure effectively. The body is descriptive rather than instructional, and the core build workflow is delegated to references rather than sequenced inline.

Suggestions

Add a brief inline numbered workflow (e.g., the first 2-3 roadmap steps) with an explicit validation checkpoint so the body stands on its own for workflow_clarity.

Include one small executable grader or eval-config snippet in the body so the main task is actionable without requiring a reference file.

DimensionReasoningScore

Conciseness

A lean reference card of compact tables that assumes Claude's competence and avoids explaining concepts Claude already knows; nearly every token earns its place.

5 / 5

Actionability

Provides a concrete tracked_metrics YAML snippet and pass@k/pass^k math, but the executable guidance for actually building evals is deferred to reference files, leaving the body's guidance incomplete on its own.

3 / 5

Workflow Clarity

The multi-step build process (roadmap steps 0-8) lives in roadmap.md rather than the body; the body itself contains no sequenced workflow or validation checkpoints.

2 / 5

Progressive Disclosure

Clear overview with well-signaled, one-level-deep references, all of which resolve to real files in ./references/, organized under a categorized Quick Reference section.

5 / 5

Total

15

/

20

Passed

Description

95%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, well-targeted description with explicit trigger guidance and a clear niche. The only minor weakness is that the stated actions are outcome-level rather than concrete techniques.

DimensionReasoningScore

Specificity

Names several concrete actions ('create evals', 'evaluate an agent', 'build evaluation suite') but they describe outcomes rather than specific techniques, leaving minor coverage gaps relative to the top anchor.

4 / 5

Completeness

Explicitly answers 'when' with concrete trigger phrases ('should be used when the user asks to...') and conveys 'what' through the action verbs, satisfying both halves clearly.

5 / 5

Trigger Term Quality

Comprehensive coverage of natural phrases users would say, including 'create evals', 'evaluate an agent', 'build evaluation suite', 'agent testing', 'graders', 'benchmarks', and agent-type variants.

5 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (evaluations for AI agents) with distinct, domain-specific triggers and minimal overlap risk with other skills.

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

Total

15

/

16

Passed

Repository
dwmkerr/claude-toolkit
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.