Invoke whenever writing, changing, reviewing, or sweeping tests in the eve repository. Authoring gate for new tests plus audit workflow for low-value, slow, implementation-coupled, or duplicative tests and the test-only production seams they demand.
69
85%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
Adapted from the OpenClaw
test-audit
skill for eve's test tiers and tooling.
Three modes, one value bar. Authoring mode gates every new or changed test at write time. Audit mode runs focused sweeps of tests that re-assert source, duplicate stronger proof, couple behavior to implementation, or keep test-only production seams alive. Continue broad audits as separate coherent follow-up PRs; optimize for confidence, not deletion count. Campaign mode prunes one whole subsystem's test surface (every test file one core area or package owns); before starting one, read CAMPAIGN.md.
Read the root AGENTS.md testing section first. It
defines eve's four tiers (unit, integration, scenario, e2e); a test belongs in
the tightest tier that can express its assertion.
Before adding any test, answer four questions; a missing answer means do not add it yet:
Then check the test against every junk pattern; a match fails the gate unless the retention bar names the contract it independently guards. A test that would break under behavior-preserving refactoring is asserting implementation, not behavior; rewrite it at the owning boundary before landing it.
Also check its cost. A scenario test that boots a dev server, installs a
tarball, or runs a real build must need that subprocess, port, or bundler. If
it only needs import "eve" to resolve or an in-memory runtime, it belongs in
integration (useTemporaryAppRoots, createTestRuntime). A test that waits
out a real production timeout or backoff should use fake timers or an existing
public option instead of wall-clock time; do not add a test-only seam to
shorten it.
Bug regression tests must fail on the pre-fix code for the intended reason and pass after the owner-boundary repair. A regression test that never demonstrably failed proves the mock, not the fix. One regression at the owner boundary covers the bug; do not replay the same scenario at every layer it crosses.
The shared checklist for both modes: the authoring gate rejects a new test that matches one, and audits hunt for existing tests that do.
Tests justify their maintenance cost by protecting behavior, a credible regression, or an independently meaningful contract. In an audit, an existing test that must change for behavior-preserving source reorganization is suspect, not automatically deletable; the authoring gate still rejects new ones.
Before judging a candidate, read the complete test and production owner, its
entry point, callers, callees, sibling implementations, overlapping tests, CI
routing (.github/workflows/), and relevant history. Read root and scoped
AGENTS.md files first. When the test claims dependency-backed behavior,
inspect the dependency source or types directly.
Keep discovery read-only and report evidence before editing. Baseline with the JSON reporter so slow files and tests are visible:
pnpm --filter eve exec vitest run --config vitest.<tier>.config.ts \
--reporter=json --outputFile=/tmp/<tier>.jsonFor broad scope, run parallel discovery lanes when available, split along production owner boundaries:
packages/eve/src/{execution,harness,runtime,compiler,...});src/internal;*.integration.test.ts);*.scenario.test.ts, packages/eve/test/scenarios/);packages/eve-code, packages/eve-catalog,
apps/docs, e2e/);Outside campaign mode, prefer a few high-confidence candidates over a large speculative inventory. Hunt for the junk patterns.
Keep a test when it independently enforces a public API (eve/* exports),
protocol, config, migration, storage, security, platform, default, prompt-byte,
package, release, or architecture contract. Also keep:
Static or slow is not a deletion reason on its own. A slow test with a unique contract gets cheaper (shared boot, fake timers, lower tier), not deleted. A test that resembles implementation may still be the independent contract; prove otherwise before removing it.
Record every field below before editing. A missing field means the candidate is not ready for deletion:
Choose one coherent owner-boundary batch. Delete obsolete test-only exports, globals, wrappers, and dead production paths instead of preserving aliases. Move retained regressions to their canonical owners. Consolidate repeated package or dependency assertions into one generic contract.
Prefer net-negative production LOC. Do not add replacement tests that restate
the same implementation, and do not convert uncertain candidates into cleanup
to increase deletion counts. Do not commit fixture trees under
packages/eve/test/fixtures/; scenario apps stay inline descriptors.
Never edit source or tests while Vitest is running in the checkout.
vitest run:
pnpm --filter eve exec vitest run --config vitest.<tier>.config.ts <path>.
Run pnpm --filter eve build:compiled first if #compiled/* changed, and
pnpm build before scenario runs.pnpm guard:invariants.pnpm fmt, pnpm lint, pnpm typecheck, then git diff --check.pnpm typecheck
and pnpm lint are the proof that nothing else depended on it.git diff --numstat; report production/tooling separately from
tests and test support. Report before and after tier timings for speed work.AGENTS.md changeset rule. Test-only changes need none;
removing dead production code from packages/eve still needs one.Commit, push, open a PR, or land only when authorized. Sign off every commit
(git commit -s) and use the
gh-pr-description skill. Land one coherent
PR at a time; after landing, refresh from current main and rerun read-only
discovery for the next high-confidence batch.
Report:
e0ad9a0
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.