CtrlK
BlogDocsLog inGet started
Tessl Logo

anthropic-evaluations

This skill should be used when the user asks to "create evals", "evaluate an agent", "build evaluation suite", or mentions agent testing, graders, or benchmarks. Also suggest when building coding agents, conversational agents, or research agents that need quality assurance.

86

1.50x
Quality

80%

Does it follow best practices?

Impact

98%

1.50x

Average score across 3 eval scenarios

SecuritybySnyk

Failed to scan

The risk profile of this skill

SKILL.md
Quality
Evals
Security

Failed to scan

The security scan for this skill could not be completed

Repository
dwmkerr/claude-toolkit

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.