Use when the user asks to find, search for, or optimize the best quantization recipe for a model, including direct requests like "find the best quantization recipe and generate a PTQ checkpoint." Guides the multi-candidate loop: choose compute-vs-memory success metrics, select ModelOpt recipe baselines, design AutoQuant/manual recipe deltas, interpret sensitivity, and decide next candidates. Do NOT use for a single known PTQ recipe run (use ptq), serving (use deployment), creating/running evals (use evaluation or launching-evals), monitoring jobs (use monitor), MLflow browsing (use accessing-mlflow), or comparing completed baseline-vs-candidate scores only (use compare-results).
80
100%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
Use this skill when quantization is an iterative recipe search, not a one-off PTQ run. The skill owns strategy: define success, choose the search space, sequence candidates, and decide the next iteration. It delegates checkpoint generation, serving, evaluation, monitoring, and metric comparison to the existing execution skills.
Treat a direct request such as "find the best quantization recipe and generate a PTQ checkpoint for this model" as enough to start. Recover local state first, then ask only for missing decisions that change the search.
ptq to produce and validate checkpoints.deployment to serve checkpoints and debug serving-specific flags.evaluation to create NEL configs and submit evals.launching-evals to run, resume, debug, and analyze NEL runs.monitor for active job tracking.accessing-mlflow for MLflow artifact lookup.compare-results for validated baseline-vs-candidate deltas and score-field comparability.Do not duplicate those workflows here. This skill should leave the user with a clear recipe portfolio, success metric, experiment sequence, and next decision.
The task is to find the best recipe for a user-defined target, not merely to produce a quantized checkpoint. A generated PTQ checkpoint is only a candidate. It becomes a recommended recipe only after evaluation and comparison against the matching baseline.
Required inputs before planning candidates:
If any of these are missing, ask for them. Do not silently default to FP8/W8A8 or call a checkpoint "best" before evaluation.
Default success rule: maximize the chosen performance objective while keeping each benchmark within 1 percentage point of the matching BF16/FP16 baseline. Near-threshold or noisy regressions require reruns before making a decision.
Keep the search space explicit. A candidate recipe is a tuple across these axes:
lm_head, adapters, vision encoders, and model-specific modules.linear_attn.in_proj_qkvz
and fused MoE expert projections such as gate/up (w1/w3).Do not collapse the search to one dimension such as numeric format only. Read
references/recipe_iteration.md when choosing concrete axes or candidates.
Recover state
monitor, launching-evals, or compare-results to recover active
job state and completed metrics when needed.Define the target
Pick baselines and first candidates
modelopt_recipes: model-specific recipes
first, then general PTQ presets or recipe fragments.Generate candidates
ptq.Gate before scaling
deployment / debug for small
patches or flags, then rerun a pipe-clean check.compare-results shows no failed external sanity check,
the candidate is comparable to the validated measured baseline, and the
user-defined goal is met. An externally unverified baseline is non-blocking.Maintain a recipe portfolio table with recipe name, objective, active-cost estimate, calibration notes, checkpoint path, eval/log references, accuracy, verbosity, positional exclusions, and decision.
references/recipe_iteration.md.references/qwen36_case_study.md only
when Qwen3.5/Qwen3.6 details are relevant.33d05b0
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.