Tune the exact WORDING of model-right-sizer's four already-shipped research-grounded layers to maximize real-execution accuracy (effort stayed within budget, per `classify_budget_adherence`) — not whether to include a layer (see `model-right-sizer-layer-ablation`), but how an included layer should be phrased. Runs a discrete coordinate-ascent search (the finite-difference analog of gradient descent for prose) over four wording knobs in `eval/tuning/knobs.py`, each anchored at one spot in the shipped agent text plausibly moving the accuracy ratio: `token_ceiling` margin, how hard the effort dial leans down under difficulty-uncertainty, and the calibration/adherence knobs. Read-mostly: never edits the agent file directly, only proposes the winning wording as a diff to review. Use when someone says "tune model-right-sizer's wording for accuracy", "optimize the budget-ceiling wording", "run a gradient descent / hill-climbing search on the agent prompt", or "which wording maximizes budget-adherence accuracy".
68
86%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Passed
No findings from the security scan
Scanned
f539a8b
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.