Invoke whenever writing, changing, reviewing, or sweeping Vitest tests in this repo. Authoring gate for new or changed tests, plus an audit workflow for low-value, implementation-coupled, or duplicative tests and the test-only production seams they keep alive. Use when asked to add tests, review tests, clean up tests, or when a PR touches *.test.ts.
75
94%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Passed
No findings from the security scan
Three modes, one value bar. Authoring mode gates every new or changed test at write time. Audit mode runs focused sweeps for tests that re-assert source, duplicate stronger proof, couple behavior to implementation, or keep test-only production seams alive. Campaign mode reviews every test file with a per-test verdict when the maintainer asks to prune the whole suite; see Campaign mode. Optimize for confidence, not deletion count. Land one coherent batch per PR; continue broad audits as separate follow-ups.
The repository's own rules live in the "Testing" section of AGENTS.md. This
skill is the procedure for applying them; when the two disagree, AGENTS.md
wins and this file needs an update.
Before adding a test, answer four questions. A missing answer means do not add it yet:
ga4-test-fixtures.ts, tool-test-support.ts) over a
near-duplicate test, and consolidate duplicated setup in the same change.src/; scripts/ counts.Then check the test against every junk pattern; a match fails the gate unless the retention bar names the contract it independently guards. A test that would break under behavior-preserving refactoring is asserting implementation, not behavior; rewrite it at the owning boundary before landing it.
Bug regression tests must fail on the pre-fix code for the intended reason and pass after the fix. A regression test that never demonstrably failed proves the mock, not the fix. One regression at the owner boundary covers the bug; do not replay it at every layer it crosses.
The shared checklist for both modes. The authoring gate rejects a new test that matches one; audits hunt for existing tests that do.
Repo-specific (each is a AGENTS.md rule):
await import() or vi.resetModules() without a comment explaining
which module-level state must reset;ga4Errors.ts, gscErrors.ts are the pattern);mockReset / mockClear ceremonies in beforeEach (Vitest clearMocks is
already on);select().from().where() with vi.fn()
chains); repositories are tested through services or real SQL;General:
Tests justify their maintenance cost by protecting behavior, a credible regression, or an independently meaningful contract. In an audit, an existing test that must change for behavior-preserving source reorganization is suspect, not automatically deletable; the authoring gate still rejects new ones.
Before judging a candidate, read the complete test and its production owner, the entry point, callers, callees, sibling implementations, overlapping tests, and relevant git history. When the test claims dependency-backed behavior (DataForSEO response shapes, Autumn, Better Auth, Drizzle), inspect the dependency source or types directly.
Keep discovery read-only and report evidence before editing. For broad scope, run parallel read-only lanes (Explore subagents work well):
src/server/mcp/ and src/server/mcp/tools/;src/server/lib/, src/server/auth/, src/server/billing/,
src/server/workflows/, src/serverFunctions/;src/server/features/;src/client/, src/shared/, src/lib/, src/types/, scripts/,
web/tests/, plus a cross-cutting grep sweep for the repo-specific patterns
above.Prefer a few high-confidence candidates over a large speculative inventory.
Use only when the maintainer explicitly asks for a suite-wide prune. Run the pipeline in two phases with disjoint file slices so agents never edit the same file:
*.test.ts into 6-8 slices balanced by
line count and grouped by directory. One reviewer per slice reads each test
file and its production owner completely and writes a report with a table
per file: test name, DELETE / MERGE / KEEP, one-line evidence naming the
code change that would fail it or the test that already owns it. The
report also lists test-only production seams (with callers checked in
src/, scripts/, tests/badseo/, web/), test support that becomes unused,
whole files to delete, and a "risky calls" section. Reviewers do not edit.Rules that hold throughout a campaign:
export keyword from a
helper whose remaining callers are all tests. Knip fails ci:check on the
orphaned export otherwise. Never remove parameters, delete branches, or
change behavior in the same change, even when a test looks like the only
reason the code exists; report those as follow-ups.src/routes/ and barrel files keep their exports.pnpm exec prettier --write <files>; pnpm format:write ignores
arguments and formats the whole repository.--sequence.shuffle.tests --sequence.seed=<n> runs. Removing reset
ceremony can surface latent order dependence; fix it by setting the mock's
default in beforeEach, never by restoring the ceremony.Keep a test when it independently enforces a public API, MCP tool contract, Zod boundary schema, config, migration, storage (both SQLite and Postgres), billing, auth, security, default, prompt-byte, or package contract. Also keep:
samSkills.test.ts
and the unique-index parity check in src/db/schema-parity.test.ts are this kind;Static or slow is not a deletion reason. A test that resembles implementation may still be the independent contract; prove otherwise before removing it.
Record every field before editing. A missing field means the candidate is not ready for deletion:
scripts/, tests/badseo/, and web/ (a profiling script is still a caller);Choose one coherent owner-boundary batch. Delete obsolete test-only exports,
wrappers, and dead production paths instead of preserving aliases (Knip will
flag survivors in pnpm ci:check). Move retained regressions to their
canonical owners. Consolidate repeated assertions into one table-driven case.
Prefer net-negative production LOC. Do not add replacement tests that restate the same implementation, and do not convert uncertain candidates into cleanup to increase deletion counts.
Never edit tests while Vitest is running in the checkout.
pnpm exec vitest run <path-or-filter>.pnpm sync-plugin-skills for plugin skill
drift).pnpm format:write, then pnpm ci:check (prettier, knip, tsc, oxlint,
plugin-skill sync). Knip failing on a now-unused export means delete the
export, not re-add a test.pnpm test.git diff --numstat; report production/tooling separately from
tests and test support.Commit, push, or open a PR only when authorized. Use the merge-ready skill
for the review and PR flow. If a verified finding exposes a recurring
invariant that .greptile/ does not capture, use maintain-greptile-rules;
do not promote one-off cleanups into permanent rules. Log repository friction
met along the way with papercuts.
Report:
db8bde1
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.