Use this skill for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected behavior. Covers both quick fixes (clear root cause) and complex debugging (unclear cause). Trigger: 'fix bug', 'fix error', 'broken', 'crash', 'doesn't work', 'fails with', 'loss NaN', 'training hangs', 'FSDP error', 'OOM'.
71
87%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
| Situation | Path |
|---|---|
| Clear error, obvious root cause, fix in <15 min | Quick Path (below) |
| Root cause unclear, multiple hypotheses | Full Protocol (Phase 1–5) |
| Distributed training issue (hang, wrong loss, sharding) | Full Protocol |
| Numerical accuracy / loss divergence | Full Protocol |
| 2+ failed fix attempts | Full Protocol |
.agents/knowledge/constraints.md for known pitfalls.pytest tests/<module>/ passes, no regressions across modalities.make quality, commit. Run /veomni-review before opening the PR or pushing a substantive update, not per commit.If not resolved in 15 min → switch to Full Protocol.
Track the phases with whatever todo/plan tool the running agent provides:
Phase 1: Investigate <symptom> -> in_progress
Phase 2: Pattern analysis -> pending
Phase 3: Hypothesis & test -> pending
Phase 4: Implement fix -> pending
Phase 5: Knowledge capture -> pending.agents/knowledge/constraints.md — many issues are known constraint violations.git log --oneline -10 — what changed recently?veomni/distributed/parallel_plan.py).veomni/distributed/sequence_parallel/).veomni/distributed/moe/).Find a working example (previous commit, different config, reference implementation).
Compare completely — diff line by line, not skim. Include config YAML, environment vars, and launcher scripts.
Identify ALL differences between working and broken code.
Check dependencies — different transformers version? Different PyTorch version?
If a package version upgrade is suspected, create isolated uv environments to bisect:
# Env A: the current default pin (the `transformers-stable` group).
uv venv .venv-a
VIRTUAL_ENV=.venv-a uv sync --active --extra gpu --dev
# Env B: the same tree with exactly one package moved.
uv venv .venv-b
VIRTUAL_ENV=.venv-b uv sync --active --extra gpu --dev
VIRTUAL_ENV=.venv-b uv pip install "<package>==<other-version>"--active is load-bearing. Without it uv sync runs in project mode and
targets .venv/, ignoring VIRTUAL_ENV — so both commands would rebuild
the main environment instead of the two you just created, which is the
opposite of what this is for. (UV_PROJECT_ENVIRONMENT works too.)
Confirm that installing the alternate version did not change other packages:
uv pip freeze --python .venv-a/bin/python > /tmp/veomni-bisect-a.freeze
uv pip freeze --python .venv-b/bin/python > /tmp/veomni-bisect-b.freeze
diff -u /tmp/veomni-bisect-a.freeze /tmp/veomni-bisect-b.freezeOnly the target package may differ. Pin or restore every non-target difference in Env B to Env A's version, then compare again before running the reproducer. If the target cannot run with that dependency set, report the compatibility conflict; a multi-package change is not a one-package bisect.
Then run the same reproducer in both envs, each with its own env
activated — the VIRTUAL_ENV= prefixes above apply only to the uv sync
lines they are attached to, not to whatever you run next:
(source .venv-a/bin/activate && <reproducer>)
(source .venv-b/bin/activate && <reproducer>)If the suspect package is transformers, you need two worktrees and two
venvs — one venv per worktree. They isolate different things and neither
substitutes for the other: a venv isolates the installed packages, a
worktree isolates the checkout. generated/ modeling lives in the
checkout, so two venvs in one worktree share a single generated/ and
regenerating it for Env B silently changes what Env A runs. Two worktrees
without separate venvs share one transformers install, which defeats the
bisect outright.
git worktree add ../bisect-a HEAD && (cd ../bisect-a && uv venv .venv && VIRTUAL_ENV=.venv uv sync --active --extra gpu --dev)
git worktree add ../bisect-b HEAD && (cd ../bisect-b && uv venv .venv && VIRTUAL_ENV=.venv uv sync --active --extra gpu --dev && VIRTUAL_ENV=.venv uv pip install "transformers==<other-version>")Regenerate generated/ inside each worktree against its own pin
(make patchgen) before running the reproducer — it is produced against
the pinned version, and a stale generated/ is itself a source of
failures.
Compare the package sets here too, using uv pip freeze --python with
../bisect-a/.venv/bin/python and ../bisect-b/.venv/bin/python, and
reconcile non-transformers dependency version differences as above. The
editable VeOmni and patchgen paths must point to their respective worktrees;
normalize those corresponding paths only when comparing the freeze output,
without changing either environment's editable installs. Run codegen and
the reproducer from each worktree with its own environment activated:
(cd ../bisect-a && source .venv/bin/activate && make patchgen && <reproducer>)
(cd ../bisect-b && source .venv/bin/activate && make patchgen && <reproducer>)Verification gate — before acting on a conclusion, check:
/veomni-review over the branch diff.Do this immediately after the fix is verified. Knowledge decays fast.
.agents/knowledge/constraints.md.agents/knowledge/architecture.md.agents/knowledge/testing.md. Only paths not
already covered by a directory-level CI entry need a new workflow line.
Use the workflow that owns the path, including the e2e workflows for
end-to-end tests; account for the GPU/NPU differences in that table.docs/ if the fix changes API behavior, config semantics, or usage patternsIf none apply, explicitly note "no new knowledge to capture."
Restart from Phase 1 if you catch yourself thinking "let me just try changing X and see", "quick fix for now, clean up later", or "it probably works, moving on".
After 3 consecutive failed fix attempts, stop fixing symptoms. Question whether the underlying approach is wrong, re-examine whether you are solving the right problem, and report the analysis to the user before continuing.
data_collator type matches the dataset.veomni/models/transformers/*/ are auto-generated — editing generated files directly will be overwritten.Include the relevant checklist when investigating.
When confidence is low or evidence is ambiguous, launch a subagent to challenge your conclusion:
You are a critical reviewer. Your job is to find flaws in the following conclusion.
## Conclusion Under Review
<the specific claim or decision>
## Evidence Presented
<the data, logs, experiments supporting the conclusion>
## Your Task
1. Does the evidence actually support the conclusion, or just correlate?
2. Generate 2+ alternative explanations consistent with the same evidence.
3. What specific observation would DISPROVE this conclusion? Has it been checked?
4. Was the experiment controlled (one variable changed at a time)?
## Output
Verdict: CONFIRMED / CHALLENGED / INSUFFICIENT_EVIDENCE
Findings: [issues found, counter-hypotheses, missing evidence]b8a3edc
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.