CtrlK
BlogDocsLog inGet started
Tessl Logo

cekura-eval-design

Use when the user asks to "generate (test) scenarios", "generate evaluators", "create an evaluator", "create evals", "create a scenario", "write a test scenario", "design a test case", "test my agent", "build eval coverage", "plan a test suite", "create red team tests", "set up test profiles", "configure conditional actions", "build a deterministic test", "design an IVR test", "write a unit test for a voice agent", "build a regression test", "scripted scenario", "structured evaluator", or "run evals". Also for CHANGING existing evaluators — "update an evaluator", "improve my evals", "make these evaluators stricter", "add a DTMF step", "fix the expected outcome", "attach metrics to these" — and for debugging how the testing agent speaks: "why did it read the number as a word", "make it spell digits", "wrong language". Covers evaluator design and review, coverage, test profiles, mock-tool data, conditional actions, and red-team / edge-case practice.

69

Quality

84%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An unusually disciplined authoring playbook: novel, executable, and rich in validation loops, with excellent workflow clarity and actionability. Its weaknesses are verbosity (repeated rules and rhetorical padding in a ~63 KB body) and progressive-disclosure discipline — heavy inline detail duplicating reference files, plus example-file references that do not exist in the bundle.

Suggestions

Trim sentence-level rhetorical justifications and consolidate rules stated in multiple sections (mode/write-path, <hold> vs <silence>) into a single statement plus a pointer, targeting roughly half the current body length while keeping every normative rule.

Move the full CA tag table and expected-outcomes rulebook into the existing references/conditional-actions.md and references/expected-outcomes.md, leaving only the decision-relevant subset inline as the write-path overview.

Create the referenced examples/ files (workflow-eval.md, red-team-eval.md, csv-eval-creation.md) or remove the dangling references from the 'Reference files (load on demand)' section.

DimensionReasoningScore

Conciseness

The body is almost entirely novel, Cekura-specific operational rules rather than explanations of concepts Claude already knows, so most tokens earn their place. However, at ~63 KB it is padded with long rhetorical justifications ("a suite that silently gains a second copy of a test is one nobody can read later", "an ungrounded assertion does not fail loudly — the condition never fires") and repeats the same rules across sections (the mode/write-path decision restated in 'Mode and write path', 'Behavioral scenarios', the self-checks, and 'Batch routing'; the <hold>-vs-<silence> distinction explained in the tags table, the tag section, 'Batch routing', and 'Personality'), which could be tightened or consolidated into references.

3 / 5

Actionability

Guidance is fully concrete and executable: a complete CA JSON payload, exact field tables with ranges and defaults (<volume ratio="1.5"> is 0–2.0, <noise time="1100"> in bare milliseconds), a per-request mode table, exact trigger phrases for mode switching, numbered refuse-to-send self-checks, and explicit poll/stall/freeze handling with real thresholds. Everything an agent needs to act is specified with no pseudocode.

5 / 5

Workflow Clarity

A clear 7-step workflow is sequenced up front, and validation is explicit at every risky point: mandatory pre-reads, a one-consolidated-checkpoint rule, self-checks before every write, post-generation reconciliation against write responses, stall/freeze retry-once loops with escalation, a 3–5 evaluator smoke cohort before large voice batches, and honest reporting that separates 'created' from 'validated'. Error-recovery feedback loops (rejected writes, blocked outcome lines, setup errors) are all spelled out.

5 / 5

Progressive Disclosure

References are one level deep and well signaled (bold in-body pointers plus a 'Reference files (load on demand)' section), but the main file carries large inline blocks — the full 20-row CA tag table, the entire expected-outcomes rulebook, test-data approaches A/B/C — that overlap the 89 KB references/conditional-actions.md and belong in the reference files. More concretely, the body points to examples/workflow-eval.md, examples/red-team-eval.md and examples/csv-eval-creation.md, and no examples/ directory exists — dangling references.

3 / 5

Total

16

/

20

Passed

Description

91%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it explicitly answers both what and when with an extensive, natural set of trigger phrases covering creation, modification, and debugging. Its only gaps are a verb-light statement of capabilities and generic testing triggers (no platform name) that create minor conflict risk with ordinary unit-test requests.

DimensionReasoningScore

Specificity

The closing sentence enumerates several concrete capability areas — "evaluator design and review, coverage, test profiles, mock-tool data, conditional actions, and red-team / edge-case practice" — giving broad, specific coverage, though they are framed as topics ("Covers …") rather than the multiple concrete action verbs of the top anchor (create, review, design, debug are only implied via the quoted triggers). Not 5: the 'what' is a single verb-light topic list; not 3: it names far more than 1–2 concrete areas.

4 / 5

Completeness

Both questions are answered explicitly: 'what' via "Covers evaluator design and review, coverage, test profiles, mock-tool data, conditional actions, and red-team / edge-case practice", and 'when' via the opening "Use when the user asks to …" with dozens of concrete quoted trigger phrases spanning create, change, and debug requests. Matches the top anchor exactly.

5 / 5

Trigger Term Quality

Comprehensive natural-language triggers including synonyms and phrasings users would actually say: "generate evaluators", "create evals", "test my agent", "build a regression test", "add a DTMF step", "make these evaluators stricter", even debugging complaints like "why did it read the number as a word" and "wrong language". Coverage of creation, modification, and debugging phrasings is exhaustive.

5 / 5

Distinctiveness Conflict Risk

The niche is fairly clear (evaluator/test-scenario design for voice/chat agents, with DTMF, IVR, and red-team specificity) and modification/debug triggers are distinctive. Not 5: the description never names the platform (Cekura), and generic triggers like "run evals", "design a test case", "build a regression test", or "write a unit test for a voice agent" would plausibly fire for ordinary unit-testing or eval-harness requests outside this platform — minor but real overlap risk with closely related testing skills.

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
cekura-ai/cekura-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.