Adversarial dynamic e2e QA workflow - generate hostile scenarios, test, verify, fix, report, and clean up
59
74%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Passed
No findings from the security scan
Fix and improve this skill with Tessl
tessl review fix ./plugins/oh-my-codex/skills/ultraqa/SKILL.mdThe canonical home for this skill is ultraqa in Yeachan-Heo/oh-my-codex
Use this explicit opt-in when a runnable behavior needs adversarial dynamic end-to-end
QA. Shared operating invariants live in templates/AGENTS.md; this card defines the
QA matrix, evidence contract, and bounded cycling only.
/ultraqa --tests|--build|--lint|--typecheck|--interactive or /ultraqa --custom "pattern" for the corresponding goal; without a structured goal, derive a runnable behavior goal.continue on the current verified next step.continue, advance the current verified QA step rather than restarting discovery.Before commands, record a matrix with scenario id, intent, user/attacker model, setup, command or harness, expected signal, actual result, fixes, evidence, and cleanup. Include a normal path and relevant hostile classes:
continue, stop/cancel/abort wording, partial output, and retries.--tests runs project tests; --build runs build plus built-artifact probes; --lint runs lint; --typecheck runs typecheck plus typed harnesses; --custom verifies pattern and exit status; --interactive uses a bounded CLI/service harness.Generate temporary tests, scripts, fixtures, or harnesses only when useful. Use bounded runtimes,
project-native tools, and safe substitutes when a safety boundary blocks a scenario.
Use absolute repo imports and pathToFileURL(join(repoRoot, "dist", ...)).href; Never rely on ./dist from /tmp.
Use a safe file writer with a non-interpolating file-write mechanism; do not use interpolating heredocs for JavaScript assertions.
Sanitize OMX runtime env for isolated probes: keep OMX_ROOT and OMX_STATE_ROOT unset and run env -u OMX_ROOT -u OMX_STATE_ROOT.
Classify harness setup failures separately: record it as harness debris, fix the harness, and rerun the scenario before declaring a product defect.
No destructive commands, secret exfiltration, credential dumping, production writes, or unbounded process spawning. Use no unbounded waits; preserve unrelated dirty work. If a scenario is unsafe, record it blocked and the safe substitute. Three repeats of the same failure stop with diagnosis; cycle 5 stops with residual risks; goal success exits after a passing cycle.
Use CLI-first lifecycle state and exact commands:
omx state write --input '{"mode":"ultraqa","active":true,"current_phase":"planning","iteration":1,"started_at":"<now>","scenario_matrix":[]}' --json
omx state write --input '{"mode":"ultraqa","current_phase":"qa","iteration":<cycle>,"scenario_matrix":"<updated matrix path or summary>"}' --json
omx state write --input '{"mode":"ultraqa","current_phase":"adversarial-e2e"}' --json
omx state write --input '{"mode":"ultraqa","current_phase":"diagnose"}' --json
omx state write --input '{"mode":"ultraqa","current_phase":"fix"}' --json
omx state write --input '{"mode":"ultraqa","current_phase":"cleanup"}' --json
omx state write --input '{"mode":"ultraqa","active":false,"current_phase":"complete","completed_at":"<now>"}' --json
omx state read --input '{"mode":"ultraqa"}' --json
omx state clear --input '{"mode":"ultraqa"}' --jsonOn completion, max cycles, same failure, safety boundary, or environment error, clean state and temporary artifacts. Report cleanup status and clean temporary e2e harnesses. Never claim complete without current evidence.
Return # UltraQA Report with: Goal and success criteria (including stop condition
and safety bounds); Scenario matrix (all columns above); Commands run (exit code,
purpose, timeout, key output); Failures found (root/user/safety impact); Fixes
applied / Fixes applied (files, rationale, scenarios, regression evidence); Cleanup and rollback
(artifacts/processes/worktree before/after); Residual risks; and Evidence
(logs, harness output, screenshots/transcripts where relevant, rerun/flake evidence).
ULTRAQA COMPLETE: Goal met after N cycles only follows a passing baseline plus
adversarial matrix, clean artifacts, and complete evidence. Otherwise return the exact
bounded status: ULTRAQA STOPPED: Max cycles, ULTRAQA STOPPED: Same failure detected 3 times,
ULTRAQA BLOCKED: ..., or ULTRAQA ERROR: ... with owner and next safe step.
1dcf513
Canonical home
since Sep 28, 2026
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.