CtrlK
BlogDocsLog inGet started
Tessl Logo

cekura-metric-design

Use when the user asks to "create a metric", "write a metric", "design a metric", "build a metric for", "evaluate agent performance", "measure call quality", "track a KPI", "add a workflow metric", "improve my metric", "fix a metric", "debug metric results", "set up quality scoring", or "what metrics do I need". Also relevant when discussing LLM judge prompts, custom code metrics, evaluation triggers, VALID_SKIP patterns, section extraction, or metric best practices for Cekura voice AI agents. Covers both creating new metrics and reviewing, iterating on, or troubleshooting existing ones.

83

1.38x
Quality

86%

Does it follow best practices?

Impact

65%

1.38x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, domain-rich playbook: the workflow is explicitly sequenced with validation and cost-guard checkpoints, and heavy detail is correctly offloaded to one-level-deep reference files. The two real defects are trimmable plumbing/redundancy in the body and the four advertised example files missing from the bundle.

Suggestions

Add the four missing files under examples/ (llm-judge-metric.md, narrative-metric.md, custom-code-metric.py, section-extraction-metric.py) or remove the 'Example Files' section — the dangling pointers break navigation to the most concrete guidance and are the main reason actionability and progressive disclosure fall short of 5.

Consolidate the three closing sections ('Next Steps', 'Documentation', 'Additional Resources') into one, and collapse the four scattered 'See references/advanced-patterns.md' pointers into a single index entry, to cut redundant tokens.

Move the verification-tag instructions and 'Performing Platform Actions' boilerplate into a short pinned preamble or reference file so the body opens with the Purpose and workflow content.

DimensionReasoningScore

Conciseness

The body is dense and almost entirely platform-specific knowledge Claude cannot know elsewhere (field-name gotchas, 'description' vs 'prompt', VALID_SKIP, cost guard), with concepts taught via compact worked examples (the spirit-vs-letter one/two-question example). It is not a 5 because of trimmable material: the verification-tag preamble and 'Performing Platform Actions' plumbing, four separate pointers to advanced-patterns.md, and overlapping closing sections ('Next Steps', 'Documentation', 'Additional Resources') that restate the same file list.

4 / 5

Actionability

Guidance is highly concrete and executable: exact API parameters ('page_size=1', 'timestamp__gte/lte'), exact field names, a fill-in trigger prompt template, deprecated types with their failure mode ('API returns 400'), and a precise manual-fix loop (categorize, patch, re-evaluate 20-30 calls). It stops short of 5 because the complete end-to-end metric examples are delegated to the four `examples/` files, which are absent from the bundle — a reader following those pointers finds nothing.

4 / 5

Workflow Clarity

The 6-step creation workflow is clearly sequenced with explicit validation ('run on sample conversations, compare to expected outcomes') and a built-in feedback loop ('Iterate... until the metric matches expectations on all samples. Plan for at least one iteration'). The batch operation is properly guarded: query call count first, stop and ask if >100, and the Manual Fix First loop re-evaluates 20-30 calls to validate each fix before escalating. This matches the top anchor's sequence-plus-explicit-validation-plus-recovery pattern.

5 / 5

Progressive Disclosure

SKILL.md is a genuine overview: templates and deep patterns are pushed to four real one-level-deep reference files (verified present and substantive), signaled inline at point of use and indexed in 'Additional Resources'. It is not a 5 because the 'Example Files' section lists four `examples/` paths (llm-judge-metric.md, narrative-metric.md, custom-code-metric.py, section-extraction-metric.py) that do not exist in the bundle — dangling references that break navigation to the most concrete material, more than a minor organization gap but leaving the overall structure good.

4 / 5

Total

17

/

20

Passed

Description

91%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with exhaustive, natural trigger-term coverage and an explicit 'Use when...' clause paired with a stated capability scope. Its only weaknesses are that the deliverable (metric prompts/code) is implied rather than named, and the improvement/debug triggers overlap with a sibling skill's territory.

DimensionReasoningScore

Specificity

The description lists several concrete actions — 'creating new metrics', 'reviewing, iterating on, or troubleshooting existing ones', 'evaluate agent performance', 'measure call quality', 'set up quality scoring' — anchored to a concrete domain (Cekura voice AI agent call-quality metrics). It stops short of a 5 because it never states what the skill actually produces (LLM judge prompts, custom_code metrics, evaluation triggers are named only as discussion topics, not deliverables), leaving minor gaps in coverage of the capability surface.

4 / 5

Completeness

Both 'what' and 'when' are explicit: 'Covers both creating new metrics and reviewing, iterating on, or troubleshooting existing ones' answers what, and an extensive 'Use when the user asks to...' clause with concrete quoted trigger phrases answers when. It is not a 4 because the 'when' is not merely present but exhaustive (13 concrete triggers plus an 'Also relevant when discussing...' clause), matching the top anchor's pattern of explicit what + concrete trigger phrases.

5 / 5

Trigger Term Quality

Thirteen natural trigger phrases cover the full verb space users would actually say ('create a metric', 'write a metric', 'design a metric', 'build a metric for', 'improve my metric', 'fix a metric', 'debug metric results', 'track a KPI', 'what metrics do I need') plus technical synonyms ('LLM judge prompts', 'custom code metrics', 'VALID_SKIP patterns', 'section extraction'). Comprehensive coverage including synonyms and question-form phrasings; not below 5 because no common variation is missing.

5 / 5

Distinctiveness Conflict Risk

'Cekura voice AI agents' establishes a clear niche with metric-specific triggers ('create a metric', 'track a KPI') that are unlikely to fire for unrelated skills. It is not a 5 because the description explicitly claims improvement territory ('improve my metric', 'fix a metric', 'debug metric results', 'troubleshooting existing ones') that overlaps with the sibling cekura-metric-improvement skill referenced in the body, creating minor cross-trigger risk between two closely related skills.

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
cekura-ai/cekura-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.