Use when the user asks to "create a metric", "write a metric", "design a metric", "build a metric for", "evaluate agent performance", "measure call quality", "track a KPI", "add a workflow metric", "improve my metric", "fix a metric", "debug metric results", "set up quality scoring", or "what metrics do I need". Also relevant when discussing LLM judge prompts, custom code metrics, evaluation triggers, VALID_SKIP patterns, section extraction, or metric best practices for Cekura voice AI agents. Covers both creating new metrics and reviewing, iterating on, or troubleshooting existing ones.
83
85%
Does it follow best practices?
Impact
100%
1.47xAverage score across 2 eval scenarios
Low
Low-risk findings worth noting
LLM judge metric design for voice AI booking flow compliance
llm_judge metric type
100%
100%
Prompt in description field
100%
100%
SCOPE & FOCUS section
100%
100%
DO NOT FLAG section
100%
100%
INPUTS — relevant variables only
100%
100%
Numbered evaluation sections
100%
100%
Explicit PASS and FAIL examples
70%
100%
FAILURE CONDITIONS — closed list
100%
100%
SAFEGUARDING NOTES
100%
100%
MM:SS timestamp requirement
100%
100%
N/A conditions covered
100%
100%
Concept-based scoping language
80%
100%
Conditional trigger and baseline metrics setup
TRUE-if-ANY phrasing
100%
100%
Do-NOT phrasing
100%
100%
Short-call exclusion
0%
100%
Expected Outcome listed
0%
100%
Infrastructure Issues listed
0%
100%
Tool Call Success listed
0%
100%
Latency listed
100%
100%
Two-step activation step 1
40%
100%
Two-step activation step 2
40%
100%
Expected Outcome audio limitation
0%
100%
Two-layer N/A strategy
60%
100%
caa6544
Table of Contents
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.