CtrlK
BlogDocsLog inGet started
Tessl Logo

cekura-metric-improvement

Use when the user asks to "improve a metric", "run labs", "leave feedback on a metric", "add to labs", "fix metric accuracy", "review metric results", "find misaligned metrics", or "iterate on metric quality". Covers the metric improvement cycle, the feedback workflow, and the labs pipeline used to refine metric accuracy over time.

66

Quality

79%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./cekura/skills/cekura-metric-improvement/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, actionable playbook with strong validation and cost-guard checkpoints. Its main weakness is redundancy: the labs workflow is described three times, and the API reference is inlined rather than split out.

Suggestions

Collapse the 'Labs Improvement Cycle' overview, the 'Guided Approach (Simulate Labs)', and the 'Interactive Labs Simulation' into a single canonical workflow to remove triplication and improve conciseness.

Move the 'API Endpoints Reference' table and its JSON payload examples into a references file (e.g. references/api-endpoints.md), keeping only the most-used endpoints inline with a clear link.

Tighten 'Manual Fix First, Then Labs' so it reads as an explicit precondition branch of the main cycle rather than a parallel workflow, reducing reader confusion about which path to follow.

DimensionReasoningScore

Conciseness

Largely avoids explaining concepts Claude already knows, but the core labs workflow is restated three times (Labs Improvement Cycle overview, Step 1-6 detail, and Interactive Labs Simulation), and the 'Guided Approach (Simulate Labs)' further overlaps Step 1, so it could be meaningfully tightened.

3 / 5

Actionability

Provides concrete HTTP endpoints with methods, JSON payloads (process_feedbacks, create_from_call_log), and specific numeric guidance ('6+ feedback instances', '20-30 calls per metric', 'page_size=1'), but some steps defer to 'see API Endpoints Reference' and rely on tool/session availability.

4 / 5

Workflow Clarity

Steps are clearly sequenced with explicit validation checkpoints (re-evaluate sample, validate changes), a cost guard requiring approval before >100 calls, and a feedback loop on validation failure; the redundant restating of the workflow across three sections slightly muddies the single canonical path.

4 / 5

Progressive Disclosure

Good section structure with a clearly signaled one-level-deep reference (references/feedback-examples.md, verified to exist); the inlined API Endpoints Reference table with JSON payloads is content that could arguably live in a separate reference file, a minor organization gap.

4 / 5

Total

15

/

20

Passed

Description

91%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with explicit what/when structure, rich natural trigger terms, and clear domain scoping. Its only weakness is mild overlap with closely related sibling metric skills.

DimensionReasoningScore

Specificity

Names the domain and several concrete actions via trigger phrases ('improve a metric', 'run labs', 'leave feedback on a metric', 'fix metric accuracy') and summarizes three concrete workflows (improvement cycle, feedback workflow, labs pipeline); minor coverage gaps keep it just below a 5.

4 / 5

Completeness

Explicitly answers both what ('Covers the metric improvement cycle, the feedback workflow, and the labs pipeline...') and when ('Use when the user asks to...') with concrete trigger phrases, matching the top anchor.

5 / 5

Trigger Term Quality

Comprehensive natural-term coverage with synonyms across the same intent ('improve a metric', 'fix metric accuracy', 'iterate on metric quality', 'review metric results', 'find misaligned metrics') plus labs-specific phrases users would naturally say.

5 / 5

Distinctiveness Conflict Risk

Labs/feedback triggers are a fairly distinct niche, but sibling skills (cekura-metric-design, cekura-eval-design) create minor overlap risk on shared terms like 'fix metric accuracy' and 'review metric results'.

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
cekura-ai/cekura-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.