CtrlK
BlogDocsLog inGet started
Tessl Logo

ce-retune

Retune a skill corpus for a new model, measurement-first: mine the run archive for a baseline, establish a noise floor, audit the corpus adversarially, then cut in measured passes until a pre-registered bar clears. Requires a benchmark harness that can A/B two builds of the corpus; refuses without one.

66

Quality

80%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/ce-retune/SKILL.md

The canonical home for this skill is ce-retune in EveryInc/compound-engineering-plugin

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a strong example of progressive disclosure: a lean overview of phases with hard gates and validation checkpoints, deferring all procedure to real, clearly signaled reference files. The only gap is the absence of concrete example artifacts (a sample registered bar, a finding format) in the body itself.

DimensionReasoningScore

Conciseness

The ~40-line body carries only operational doctrine and hard rules — it explains nothing Claude already knows and defers all procedure to the reference files — so every token earns its place, matching the lean-and-efficient anchor rather than the minor-trimmings anchor below.

5 / 5

Actionability

Instructions are concrete and decisive ("run the harness against two identical copies of the corpus, same commit on both sides", "Register the bar now, in writing, before any change exists", "One agent per skill proposes cuts; a second per skill does the opposite"), but the body includes no concrete example artifacts such as a sample registered bar or scoring table, leaving minor gaps that keep it below fully-executable copy-paste-ready guidance.

4 / 5

Workflow Clarity

A clear sequence (Phase 0 gate then phases 1-6) with explicit validation steps (the measurement gate, the pre-registered bar, "stop and say so"), an error-recovery feedback loop ("Loop 4 and 5 until the registered bar clears", "measure, then let the failure choose the next fix"), and a Phase-0 checklist — matching the top anchor even though the operations are destructive/batch, since validation is thorough.

5 / 5

Progressive Disclosure

The body is a clean overview with six well-signaled, one-level-deep references, each phase naming the reference it cannot start without; all referenced paths resolve to real files in references/ with no deeper nesting, matching the clear-overview anchor.

5 / 5

Total

19

/

20

Passed

Description

67%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific, method-rich, and clearly distinct from other skills, with a strong what. Its weaknesses are the absence of an explicit trigger clause and reliance on measurement jargon over the natural phrases a user would say when a corpus degrades on a new model.

Suggestions

Add an explicit trigger clause, e.g., "Use when skills degrade or regress after a model switch or upgrade, or when preparing a skill corpus for a new model", to lift completeness past the cap.

Include natural user phrasings and synonyms alongside the jargon — "skill performance dropped", "corpus regressed", "model upgrade/switch" — so trigger_term_quality reflects how users actually describe the problem.

Keep the A/B-harness requirement (it drives distinctiveness) but phrase it in user-facing terms such as "needs a runner that can A/B two corpus builds" so a user can recognize whether they have one.

DimensionReasoningScore

Specificity

The description lists multiple concrete, distinct actions forming a complete method — "mine the run archive for a baseline, establish a noise floor, audit the corpus adversarially, then cut in measured passes until a pre-registered bar clears" — which matches the comprehensive-coverage anchor; there are no coverage gaps that would place it at 4.

5 / 5

Completeness

The "what" is explicit and detailed, but there is no "Use when..." clause or equivalent explicit trigger guidance; "for a new model" only weakly implies when to invoke the skill, so per the judging guideline the missing trigger clause caps completeness at 3.

3 / 5

Trigger Term Quality

Relevant domain keywords are present ("Retune a skill corpus", "noise floor", "benchmark harness"), but they are measurement/internal jargon and the natural user phrasings and synonyms (e.g., "skills degraded after a model switch", "skill performance regressed", "model upgrade") are missing, matching the "some relevant keywords but missing common variations" anchor rather than the good-coverage anchor above.

3 / 5

Distinctiveness Conflict Risk

A clear niche with a hard precondition ("Requires a benchmark harness that can A/B two builds of the corpus; refuses without one") makes it highly distinguishable from other skills with minimal conflict risk, matching the top anchor.

5 / 5

Total

16

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
crdant/compound-engineering-plugin
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.