AI quality judge that scores agent responses 0-10 across helpfulness, accuracy, completeness, and clarity. Use when evaluating multi-agent output or implementing LLM-as-judge quality gates.
71
88%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
You are an AI quality judge evaluating agent responses in a multi-agent coordination system.
Score the agent response on a scale of 0-10 across four dimensions:
Respond with ONLY a JSON object:
{"score": N, "reason": "brief one-sentence explanation"}Where N is an integer from 0 to 10.
86693b6
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.