Land the winning experiment from an evo run as a clean, mergeable change -- open a PR when the repo has a remote, otherwise merge into the working branch. Distills the best-scoring experiment down to the minimal diff that reproduces its behaviour, shaped for the qualities a maintainer merges on (scope discipline, test integrity, style adherence), then attaches an advisory mergeability report. Use when the user invokes /evo:ship, asks to land/merge/ship the best result, or wants to turn a finished optimization into a pull request.
77
98%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
Turn a finished evo run into a change a maintainer would merge.
The optimize loop leaves a tree of committed experiments. The winning worktree diff is not mergeable as-is: it carries debug prints, search-process churn, over-broad edits, and sometimes a test that was relaxed to clear a gate. Shipping is the step that re-derives the minimal clean change reproducing the winning behaviour, lands it the way the repo expects (PR or merge), and reports how mergeable it is.
Correctness is the floor, not the goal. The score says the behaviour works; this skill decides whether the diff is fit to merge.
/evo:ship # ship the auto-selected winner
/evo:ship exp_0042 # ship a specific experiment insteadPick the experiment to ship, then confirm it with the user before touching their tree.
evo status # current best valid score + counts
evo report # top valid experiments table + score chartevo frontier is for choosing where to branch next; it can
exclude an exhausted branch whose score is still the right thing to ship. An
explicit exp_id argument overrides auto-selection.committed, or pruned with
prune_kind=exhausted, with a commit and score, no gate_result === false,
and no invalid-pruned ancestor. Never select discarded, failed, active,
evaluated, legacy-pruned nodes with no prune_kind, prune_kind=invalid, or
descendants of invalid-pruned nodes. If no valid candidate exists, stop and
report why nothing is safe to ship.evo diff <root_id> <winner_id> # target-scoped cumulative diff, baseline -> winnergit diff <baseline_commit> <winner_commit>); each node carries .commit.Work on a fresh branch off the user's current HEAD, not in the experiment worktree. Re-derive the change so it stands on its own:
Then confirm the behaviour survived the distillation:
evo run <winner_id> --check # or the project's benchmark / test commandIf the distilled change no longer reproduces the winning score, do not paper over it -- report the gap (which part of the experiment diff was load-bearing) and let the user decide. Best-effort means honest about what could not be cleaned up, not silently shipping the raw worktree.
Detect how the repo expects changes to arrive:
git remote -vgh pr create with the mergeability report (Stage 4) as the
body. Do not push or open the PR without the user's go.The landed commit message carries provenance: the winning experiment id, the score delta, and the one-line hypothesis. State what changed and why it is safe; do not narrate the search process.
Always produce the report. It never blocks the merge -- it tells the user, and a future reviewer, how mergeable the change is across the axes a maintainer judges on:
Lead with a plain-language summary: what changed and why it is safe to merge. On a remote repo this is the PR body. With no remote, print it and save it alongside the run so the user can paste it into a review later.
Everything above is method you can adapt to the repo. These are not:
ab5fbd6
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.