CtrlK
BlogDocsLog inGet started
Tessl Logo

cekura-metric-improvement

Use when the user asks to "improve a metric", "run labs", "leave feedback on a metric", "add to labs", "fix metric accuracy", "review metric results", "find misaligned metrics", or "iterate on metric quality". Covers the metric improvement cycle, the feedback workflow, and the labs pipeline used to refine metric accuracy over time.

70

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-organized operational playbook: the workflow is explicitly sequenced with validation and feedback loops, batch operations are gated by a cost guard, and guidance is concrete down to endpoint payloads and gotcha notes. The main improvement areas are the unauthenticated API fallback path and moving bulk endpoint details out of the main file.

Suggestions

Add authentication instructions (API key/OAuth setup) for the direct-API fallback path mentioned in 'Performing Platform Actions', so the skill is executable when platform tools are absent.

Move the API Endpoints Reference table and JSON payload examples into a references file (e.g., references/api-endpoints.md), keeping only the 2–3 most-used endpoints inline, to tighten progressive disclosure and conciseness.

Trim the Purpose section and the redundant checklists in Steps 3–4 to reduce token overhead without losing information.

DimensionReasoningScore

Conciseness

The body is dense and operational — endpoint tables, exact JSON payloads, parameter notes like "metrics must be an array of objects... Passing bare IDs returns 500" — with no tutorials on concepts Claude already knows. Minor trimming is possible (the Purpose section restates the description, and some checklist bullets in Steps 3–4 could be condensed).

4 / 5

Actionability

Mostly executable: concrete endpoints with full JSON bodies, pagination and filter parameters, and a cost guard with a specific procedure. The gap is the fallback path — it says to fall back to "direct API endpoints or dashboard guidance only when no tools are available" but never explains how to authenticate to those endpoints, so that branch is not fully executable.

4 / 5

Workflow Clarity

The six-step labs cycle is clearly sequenced with explicit validation checkpoints: Step 5 re-runs improved metrics and checks for regressions, "If validation fails, leave additional feedback and iterate" provides a feedback loop, and the cost guard requires counting and user confirmation before any bulk operation — satisfying the validation requirement for batch operations.

5 / 5

Progressive Disclosure

Good structure with clear section headers; the one bundle file (references/feedback-examples.md) exists and is properly signaled one level deep, and detailed feedback examples are correctly deferred to it. Minor gap: the ~50-line inline API Endpoints Reference with JSON payload examples is arguably bulk detail that could live in a reference file, though its use across every step keeps it defensible inline.

4 / 5

Total

17

/

20

Passed

Description

91%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: explicit 'Use when' trigger guidance with abundant natural phrases, a clear statement of scope, and a well-differentiated niche. Its only weakness is that the 'what' is expressed as covered topic areas rather than sharp concrete actions.

DimensionReasoningScore

Specificity

The description names the domain ("metric improvement cycle, the feedback workflow, and the labs pipeline") and lists several concrete capability areas implied by its triggers ("leave feedback on a metric", "fix metric accuracy", "run labs"), with minor gaps — the 'what' is stated as covered topics rather than explicit verb-on-object actions.

4 / 5

Completeness

Explicitly answers 'what' ("Covers the metric improvement cycle, the feedback workflow, and the labs pipeline used to refine metric accuracy over time") and 'when' via a leading "Use when the user asks to..." clause with concrete trigger phrases — both are present and specific.

5 / 5

Trigger Term Quality

Eight quoted natural trigger phrases with synonym coverage across both the action ("improve", "fix", "iterate on", "review") and object ("metric", "metric accuracy", "metric quality", "metric results") dimensions — a user needing this skill would plausibly say one of these.

5 / 5

Distinctiveness Conflict Risk

The Cekura labs/metric-improvement niche is well-carved and unlikely to trigger for unrelated skills, but phrases like "review metric results" and "improve a metric" overlap with closely related skills in the same family (e.g., cekura-metric-design) — mostly distinct with minor overlap risk.

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
cekura-ai/cekura-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.