CtrlK
BlogDocsLog inGet started
Tessl Logo

model-right-sizer-holdout-tuning

Tune model-right-sizer's wording knobs (`eval/tuning/knobs.py`) against a REAL, already-measured build's actual token spend — not the synthetic benchmark `model-right-sizer-prompt-tuning` searches, nor a fresh build per candidate. Picks a task from `overfitting_guard.py`'s `HOLDOUT_TASKS` registry (real actual/budgeted pairs already recorded), dispatches 3 INDEPENDENT BLIND dry-runs per candidate (no calibration-ledger access), averages the budget across draws (single-draw noise can flip within/over-budget calls), maps to real actuals, scores via `classify_budget_adherence` + `score_candidate`, diagnoses the miss pattern, proposes ONE wording change, and re-runs to check improvement. n stays fixed per task — flags rather than silently continues once squeezing looks like overfitting. Use when someone says "tune the knobs against this blueprint/build", "iterate the dry run with no prior context against the real actuals", "keep tuning until N%", or "how close does a blind estimate get to the actual cost".

68

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

No evaluations available

This skill hasn't been evaluated yet

Log in to request
Repository
Cloudzero/cloudzero-claude-marketplace

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.