CtrlK
BlogDocsLog inGet started
Tessl Logo

model-right-sizer-prompt-tuning

Tune the exact WORDING of model-right-sizer's four already-shipped research-grounded layers to maximize real-execution accuracy (effort stayed within budget, per `classify_budget_adherence`) — not whether to include a layer (see `model-right-sizer-layer-ablation`), but how an included layer should be phrased. Runs a discrete coordinate-ascent search (the finite-difference analog of gradient descent for prose) over four wording knobs in `eval/tuning/knobs.py`, each anchored at one spot in the shipped agent text plausibly moving the accuracy ratio: `token_ceiling` margin, how hard the effort dial leans down under difficulty-uncertainty, and the calibration/adherence knobs. Read-mostly: never edits the agent file directly, only proposes the winning wording as a diff to review. Use when someone says "tune model-right-sizer's wording for accuracy", "optimize the budget-ceiling wording", "run a gradient descent / hill-climbing search on the agent prompt", or "which wording maximizes budget-adherence accuracy".

68

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, disciplined runbook: explicit costs, gated spend, an independent-evaluation protocol with feedback loops for stall recovery, and honest held-out/overfit reporting make the workflow itself exemplary. Weaknesses are moderate: copy-paste command invocations are missing, and the skill is structurally dependent on DESIGN.md and the eval/tuning scripts that are not shipped inside the bundle, leaving the body's "read that first" instruction unverifiable from the skill alone. Minor redundancy (repeated never-edit rule, duplicate link lists) costs a little conciseness.

Suggestions

Progressive disclosure: ship the load-bearing dependencies inside the skill bundle (e.g. references/DESIGN-excerpt.md, scripts/ for knobs/optimizer/generate_variant) or inline the essential content — currently every ../../eval/tuning/* path lives outside the bundle while the body says 'read that file first' and Step 4 requires restating DESIGN.md's 'Known limitations' verbatim.

Actionability: add copy-paste-ready command lines for the tooling the loop depends on (e.g. the exact generate_variant.py and optimizer invocations for rendering and scoring a neighbor), since only function signatures are currently given.

Conciseness: remove the duplication between Step 4, 'What this skill does NOT do', and 'Related' — the never-edit rule is stated twice and all four eval/tuning files are linked inline and again in Related.

DimensionReasoningScore

Conciseness

The body is dense and almost entirely operational — concrete costs ("typically 3–9 real builds per candidate", "sonnet 38,401 / haiku 28,508 tokens"), specific tool functions, and no explanation of concepts Claude already knows. It falls at the 4 anchor ("Efficient; minor instances of over-explanation that could be trimmed") rather than 5 because of real duplication: the "never edits the shipped file" rule is stated in both Step 4 and "What this skill does NOT do", the `../../eval/tuning/` files are each linked inline and again in "Related", and asides like "this SKILL.md is the runbook; the design doc is the honest account" are editorial padding. It is clearly not 3 — no whole section is unnecessary.

4 / 5

Actionability

Guidance is mostly executable with named interfaces throughout: "`knobs.ALL_KNOBS`", "`optimizer.propose_neighbors`", "`optimizer.score_candidate(records)`", "`optimizer.coordinate_ascent_step(current_settings, current_score, knob_name, neighbor_evaluations)`", the concrete tuning/held-out task IDs, and a worked recovery reference for the stall case. It sits at 4 ("Mostly executable guidance; concrete code or commands with minor gaps") rather than 5 because there are no copy-paste-ready command lines — e.g. how to actually invoke `generate_variant.py` or `optimizer` from the CLI is never shown, so the operator must reconstruct invocations from function signatures.

4 / 5

Workflow Clarity

The process is explicitly sequenced (scope confirmation → Step 1 floors → Step 2 search loop → Step 3 held-out check → Step 4 report) with validation checkpoints throughout: a spend-confirmation gate before any builds ("ask before spending real build budget"), a convergence criterion ("if no knob moved, stop — converged"), a held-out overfit check whose failure is elevated to "the headline finding", and an explicit feedback loop for stall recovery ("diagnose via the run's `journal.jsonl` (`started` vs. `result` event counts) and output-file mtimes... recover by recomputing the missing set from the filesystem"). This matches the 5 anchor (explicit validation steps, feedback loops for error recovery, a checklist for the final report); the batch-operation cap does not apply because validation is present.

5 / 5

Progressive Disclosure

Structure and signaling are good — references are one level deep and clearly marked (DESIGN.md, knobs.py, optimizer.py, generate_variant.py, the sibling skill), and the body is well-sectioned. But the score falls to the 3 anchor rather than 4 because the runbook's load-bearing content is deferred to files that are not part of the skill bundle: no `references/`, `scripts/`, or `assets/` exist, and every "[`../../eval/tuning/...`]" path resolves outside the bundle (unresolvable in this layout), while the body insists "**read that file first**" and Step 4 requires restating "Every limitation named in DESIGN.md's 'Known limitations' section" — content the bundle does not ship. The duplicated "Related" link list adds organization noise but is secondary to the unshipped-dependency gap; it is not 2 because the SKILL.md itself is a coherent, well-organized runbook rather than an inlined dump.

3 / 5

Total

16

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An excellent frontmatter description: highly specific, third-person, with a clear what/when pair, explicit trigger phrases in the user's own words, and explicit disambiguation from the sibling layer-ablation skill. The only weakness is trigger coverage narrowly centered on this skill family's jargon, missing a few plainer synonyms a user might say.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions with comprehensive coverage: "Runs a discrete coordinate-ascent search (the finite-difference analog of gradient descent for prose) over four wording knobs", "never edits the agent file directly, only proposes the winning wording as a diff to review", and names the concrete knobs ("`token_ceiling` margin, how hard the effort dial leans down under difficulty-uncertainty, and the calibration/adherence knobs"). It clearly matches the 5 anchor ("Lists multiple specific concrete actions; comprehensive coverage") rather than 4, since no capability gaps remain — even the read-only safety boundary is stated.

5 / 5

Completeness

Both questions are answered explicitly: the 'what' ("Tune the exact WORDING of model-right-sizer's four already-shipped research-grounded layers to maximize real-execution accuracy") and the 'when' via a concrete "Use when someone says..." clause with four quoted trigger phrases. This matches the 5 anchor verbatim in structure; it is not 4 because the 'when' is explicit and multi-phrase, not merely present-but-imprecise.

5 / 5

Trigger Term Quality

Trigger phrases are natural and user-voiced: "tune model-right-sizer's wording for accuracy", "optimize the budget-ceiling wording", "run a gradient descent / hill-climbing search on the agent prompt", "which wording maximizes budget-adherence accuracy" — including the synonym pairing "gradient descent / hill-climbing". It sits at the 4 anchor ("Good keyword coverage; a few natural terms missing") rather than 5 because common variations a user might actually say, such as "improve the prompt", "make the agent more accurate", or "prompt optimization", are absent, and coverage is centered on this one sibling-skill vocabulary.

4 / 5

Distinctiveness Conflict Risk

It carves out a clear niche and actively disambiguates the closest sibling: "not whether to include a layer (see `model-right-sizer-layer-ablation`), but how an included layer should be phrased". The triggers are scoped to wording-tuning phrasing, so overlap risk with the ablation or benchmark skills is minimal — a clean fit for the 5 anchor ("Clear niche with distinct triggers; minimal conflict risk").

5 / 5

Total

19

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 13 suspicious

Warning

Total

14

/

16

Passed

Repository
Cloudzero/cloudzero-claude-marketplace
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.