CtrlK
BlogDocsLog inGet started
Tessl Logo

model-right-sizer-holdout-tuning

Tune model-right-sizer's wording knobs (`eval/tuning/knobs.py`) against a REAL, already-measured build's actual token spend — not the synthetic benchmark `model-right-sizer-prompt-tuning` searches, nor a fresh build per candidate. Picks a task from `overfitting_guard.py`'s `HOLDOUT_TASKS` registry (real actual/budgeted pairs already recorded), dispatches 3 INDEPENDENT BLIND dry-runs per candidate (no calibration-ledger access), averages the budget across draws (single-draw noise can flip within/over-budget calls), maps to real actuals, scores via `classify_budget_adherence` + `score_candidate`, diagnoses the miss pattern, proposes ONE wording change, and re-runs to check improvement. n stays fixed per task — flags rather than silently continues once squeezing looks like overfitting. Use when someone says "tune the knobs against this blueprint/build", "iterate the dry run with no prior context against the real actuals", "keep tuning until N%", or "how close does a blind estimate get to the actual cost".

68

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

73%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A rigorous, well-sequenced workflow with strong validation checkpoints and highly concrete guidance, but at ~245 lines it carries notable narrative and rationale padding that could be tightened or split into reference files without losing operational clarity.

Suggestions

Trim or move the session-history narrative (the dispatch_floor_awareness level 0→4 origin story) out of the body — keep the operational rule it motivated, cut the chronology.

Consolidate the blind-run discipline, which is currently restated in the intro, the 'Do NOT invoke' warning, step 3, and step 10, into a single authoritative statement plus a checklist reference in step 3.

Split the 'Two specific wording pitfalls' subsection into a reference file (e.g. references/wording-pitfalls.md) and keep a one-line pointer in step 7, reducing inline body length.

DimensionReasoningScore

Conciseness

The body is dense with repo-specific operational knowledge and never explains generic concepts, but it carries real over-explanation: a session-history narrative ('this exact loop is how `dispatch_floor_awareness`... went from level 0 to level 4 across several real sessions — first a novel-use-case validation exposed the gap...') and repeated restatements of the blind/no-calibration-access discipline across multiple sections. Not 4: these are whole passages of rationale and history that could be trimmed or moved, not minor instances.

3 / 5

Actionability

Concrete and specific throughout — exact scripts with flags (`generate_variant.py --settings "..."`), exact functions (`statistics.mean`, `reasoning_budget.budget_adherence_ratio`, `classify_budget_adherence`, `optimizer.score_candidate`), exact parameters (`mode: "dry_run"`, Pass A only), and the exact test to run (`tests/model_right_sizer/test_tuning_knobs.py`). Not 5: the agent-dispatch mechanics and the rendered comparison table are described but not given as executable commands or templates, leaving minor gaps.

4 / 5

Workflow Clarity

Ten clearly sequenced steps with explicit validation checkpoints (step 4 validates the blueprint before mapping, step 10 runs the full validator list and pytest suite before committing) and a built-in feedback loop (step 8 re-render/re-dispatch/re-score/compare; step 9 explicit iterate/stop/escalate decision rules). The destructive/batch cap does not apply because validation steps are present and explicit.

5 / 5

Progressive Disclosure

No bundle files exist, and the body is well-sectioned ('Before starting', 'What to do', 'What this does NOT do', 'Related') with clearly signaled one-level-deep markdown links to repo files. Not 5: substantial inline content — the 'Two specific wording pitfalls' subsection and the origin narrative — is arguably reference material that a leaner SKILL.md would point to rather than inline; not 3: structure and reference signaling are good, not merely present.

4 / 5

Total

16

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A highly specific, complete description that names concrete actions, gives explicit quoted trigger phrases, and explicitly differentiates the skill from its closest sibling. The only weakness is that the trigger phrases are jargon-leaning and could include more natural phrasing variations.

DimensionReasoningScore

Specificity

Lists multiple concrete actions end-to-end — 'Picks a task from `HOLDOUT_TASKS`', 'dispatches 3 INDEPENDENT BLIND dry-runs per candidate', 'averages the budget across draws', 'scores via `classify_budget_adherence` + `score_candidate`', 'proposes ONE wording change, and re-runs to check improvement' — with comprehensive coverage of the tuning loop. Not 4: there are no gaps in the action coverage; every phase of the workflow is explicitly named.

5 / 5

Completeness

Clearly answers both 'what' (a fully specified tuning loop against a real build's actuals) and 'when' (an explicit 'Use when someone says...' clause with concrete quoted trigger phrases). Matches the anchor for a complete what+when description with concrete triggers.

5 / 5

Trigger Term Quality

The 'Use when' clause quotes four genuinely natural user utterances ('tune the knobs against this blueprint/build', 'keep tuning until N%', 'how close does a blind estimate get to the actual cost'), giving good keyword coverage. Not 5: the phrases lean technical and miss simpler synonyms/variations a user might state the same request with; not 3: multiple natural trigger phrases are explicitly present.

4 / 5

Distinctiveness Conflict Risk

Explicitly disambiguates from its nearest neighbor — 'not the synthetic benchmark `model-right-sizer-prompt-tuning` searches, nor a fresh build per candidate' — carving out a clear niche with distinct triggers and minimal conflict risk.

5 / 5

Total

19

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 14 suspicious

Warning

Total

14

/

16

Passed

Repository
Cloudzero/cloudzero-claude-marketplace
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.