Tune model-right-sizer's wording knobs (`eval/tuning/knobs.py`) against a REAL, already-measured build's actual token spend — not the synthetic benchmark `model-right-sizer-prompt-tuning` searches, nor a fresh build per candidate. Picks a task from `overfitting_guard.py`'s `HOLDOUT_TASKS` registry (real actual/budgeted pairs already recorded), dispatches 3 INDEPENDENT BLIND dry-runs per candidate (no calibration-ledger access), averages the budget across draws (single-draw noise can flip within/over-budget calls), maps to real actuals, scores via `classify_budget_adherence` + `score_candidate`, diagnoses the miss pattern, proposes ONE wording change, and re-runs to check improvement. n stays fixed per task — flags rather than silently continues once squeezing looks like overfitting. Use when someone says "tune the knobs against this blueprint/build", "iterate the dry run with no prior context against the real actuals", "keep tuning until N%", or "how close does a blind estimate get to the actual cost".
68
85%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Passed
No findings from the security scan
Scanned
f539a8b
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.