CtrlK
BlogDocsLog inGet started
Tessl Logo

cekura-self-improving-agent

Use to close the loop on agent quality — turn a failure signal into a verified fix. Triggers: "improve my agent", "self-improving agent", "auto-tune / iterate on my prompt", "fix my agent from test results" — and production-call bug fixing: "fix this prod call issue", "debug and fix call ID", "reproduce this production bug", "regression test before raising PR". Works on any stack: provider-dashboard agents (VAPI, Retell, ElevenLabs, Bland) AND agents whose config lives in the customer's own repo, database, or prompt registry with the provider agent created at deploy time, or with custom mock servers. Fixed safety invariants (must-fail-first reproduction, attestation, no-prod-in-loop) + a per-project capability manifest declaring where config lives and how to read/render/apply/deploy/verify it.

66

Quality

78%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./cekura/skills/cekura-self-improving-agent/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

66%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-engineered orchestration overview with exceptional gating, validation, and stop-condition logic, and mostly concrete tool-level guidance. Its two real weaknesses are density (invariant 1 inlines detail that belongs in reference files) and a broken bundle: nearly all phase, recipe, and hook files it points to are missing, while three unreferenced legacy reference files sit orphaned in references/.

Suggestions

Ship or remove the dangling references: the seven `phases/*.md` files, three `recipes/*.md` files, `hooks/repro-gate.sh`, and `providers/<mode>/` playbooks are all cited in the body but absent from the bundle — the phase table's 'Re-read the phase file on entry' instruction fails on every phase.

Link or delete the orphaned `references/` files (`dynamic-variables-debugging.md`, `phase-2-failure-collection.md`, `phase-3-diagnosis.md`): they are never mentioned in the body and use an older numbered-step framing that conflicts with the v3 capability-manifest model described here.

Split invariant 1's ~40-line nested paragraph into a short invariant statement in SKILL.md plus a reference file for the reproduction-mode sizing rules, hook mechanics, and gate_override record shape; also fix the Parameters section's mid-sentence line break ('...to verify. / Explicit user overrides').

DimensionReasoningScore

Conciseness

Almost no space is wasted on concepts Claude already knows, but the body is dense to the point of being hard to parse: invariant 1 is a ~40-line nested wall of text inlining the hook mechanics, the full gate_override JSON shape, and the stochastic sizing formula, and the Parameters section has a formatting glitch ("...to verify.\nExplicit user overrides..."). Mostly efficient but several sections clearly could be tightened or pushed to reference files, which matches anchor 3 better than anchor 4.

3 / 5

Actionability

Much of the guidance is concretely executable: the exact tool call (`mcp__cekura__cekura_skill_started` with literal `skill_name`, `verification_tag`, `plugin_version` values), the exact `repro.json` field list, the run-naming format `[selfimprove:<session_id>] <phase> — <detail>`, the formula `N = clamp(⌈2/p̂⌉, 4, 10)`, and explicit numeric defaults. Not a 5 because the actual per-phase execution steps are delegated to phase files rather than shown, and no sample manifest snippet or worked command example appears in the body itself.

4 / 5

Workflow Clarity

The seven-phase table gives a clear numbered sequence with each phase's purpose, and validation checkpoints are unusually explicit: must-fail-first gate with artifact, readback attestation after every deploy, overfitting gate, budgets/stop conditions (oscillation, same failure shape 3×), "revert on any" regression damage, and a drift/failure classification table that tells the agent exactly when to stop. Feedback loops (fix → re-validate, flake retry policy, manifest repair pass) are present, matching anchor 5.

5 / 5

Progressive Disclosure

Structure and signaling are well designed — a phase table, an on-demand reference section, clearly labeled recipes — but scored against the actual bundle the navigation is largely broken: `phases/setup.md`, `phases/collect.md`, `phases/debug.md`, `phases/reproduce.md`, `phases/loop.md`, `phases/regression.md`, `phases/promote.md`, all three `recipes/*.md`, `hooks/repro-gate.sh`, and `providers/<mode>/` playbooks do not exist; only `references/manifest.schema.json` and `references/manifest-guide.md` are present. Additionally, three orphaned files in `references/` (`dynamic-variables-debugging.md`, `phase-2-failure-collection.md`, `phase-3-diagnosis.md`) are never referenced from the body and appear to be from an older numbered-step version of the skill. Since the body instructs the agent to "re-read the phase file on entry", navigation fails in practice, which sits below anchor 3's 'references present but not clearly signaled'.

2 / 5

Total

14

/

20

Passed

Description

91%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it pairs a concrete capability statement with an explicit, natural-language trigger list covering both its improve-loop and production-debug use cases, and scopes the supported stacks precisely. Slight deductions for abstract framing in the opening clause and minor trigger overlap with adjacent prompt-tuning skills.

DimensionReasoningScore

Specificity

Concrete actions and mechanisms are named — "turn a failure signal into a verified fix", "must-fail-first reproduction, attestation, no-prod-in-loop", and the manifest's "read/render/apply/deploy/verify" — plus concrete stack coverage ("VAPI, Retell, ElevenLabs, Bland... repo, database, or prompt registry... custom mock servers"). Not a 5 because the opening framing ("close the loop on agent quality") is abstract and the enumerated actions describe invariants and mechanics rather than a plainly comprehensive list of what the skill does (regression sweeps and promotion are only implied).

4 / 5

Completeness

Explicitly answers both questions: what ("turn a failure signal into a verified fix", with invariants and manifest mechanics) and when (a dedicated "Triggers:" clause listing concrete trigger phrases for both the improve loop and production-call bug fixing). This mirrors the anchor-5 example structure of capability statement plus explicit 'Use when' triggers.

5 / 5

Trigger Term Quality

Comprehensive natural-language triggers with synonyms across both use cases: "improve my agent", "self-improving agent", "auto-tune / iterate on my prompt", "fix my agent from test results", "fix this prod call issue", "debug and fix call ID", "reproduce this production bug", "regression test before raising PR" — these are phrases a user would plausibly say verbatim. Clearly matches the anchor-5 example's coverage including variants; nothing common is missing.

5 / 5

Distinctiveness Conflict Risk

The niche is clear — reproduction-gated agent fixing with safety invariants on agent stacks — and most triggers ("fix this prod call issue", "regression test before raising PR") are distinct. Minor overlap risk remains with closely related skills: "auto-tune / iterate on my prompt" could plausibly match general prompt-improvement or metric-improvement skills, which matches anchor 4 rather than the minimal-conflict anchor 5.

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
cekura-ai/cekura-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.