Concrete, per-area proof that a change actually works before reporting it fixed, done, or "should work now" — which dev server, test command, or invocation proves a template UI change, an action, a migration, a guard, or a core/package change. Use before every wrap-up, and before stopping mid-task to ask permission instead of continuing.
Before telling the user a fix, feature, or bug is done, choose the smallest proof that exercises the behavior you changed and use the row below as a guide. "I changed the code and it looks right" is not proof — it's the exact gap this skill exists to close. A failing check is a reason to keep working, not a reason to stop, ask, or report done anyway.
Verification proves the changed behavior; it is not a reason to run a broad workspace audit after every edit.
pnpm run prep, browser automation, or restart a dev server unless the
changed area requires it.pnpm run prep for shared contracts,
cross-cutting changes, migrations, or when a focused check exposes a wider
failure. Keep repository-required guards and doctor checks when they apply.| You changed | What proves it | Command |
|---|---|---|
Template UI/behavior (templates/<app>/app/**) | Drive the real page and check console + network, not just the diff | pnpm --filter <app> dev (or root pnpm dev for the gateway), then click through the exact flow with whatever browser tool is available; check for console errors and failed (4xx/5xx) requests on that page |
An action (templates/<app>/actions/*.ts) | Call it with representative args and inspect the real return value | cd templates/<app> && pnpm action <name> --key value; for a write, follow with pnpm action db-query --sql "SELECT ..." to confirm the row actually landed |
| Schema/migration | Boot the app so migrations run, then read back the new column/table | pnpm --filter <app> dev once, then cd templates/<app> && pnpm action db-query --sql "..." (action is a per-template script, not a root one); pnpm guard:additive-migrations catches destructive DDL before CI does |
A guard/lint script (scripts/guard-*.{mjs,ts}) | Run it directly against a case that should now pass and one that should still fail | pnpm guard:<name> (name matches the package.json script); pnpm guards for the full sweep |
packages/core or another publishable package | Run that package's actual tests, not just typecheck | pnpm --filter @agent-native/core exec vitest --run <changed.spec.ts>, or pnpm test:core-integration for cross-cutting paths |
| Cross-cutting change, or unsure which area | Workspace-wide pass | pnpm run prep (fmt + typecheck + test:fast + guards, run in parallel) |
| Deployed behavior (when requested or environment-specific) | Add a deployed check only when requested or a concrete beta/production difference could make local proof misleading (for example, deploy-time config or secrets, host/cookie/OAuth settings, CDN/SSR, serverless runtime, storage, or external connectivity). A merge-triggered deploy alone is not a reason to retest every path remotely; focused local proof and required CI are the default. When warranted, exercise the affected path on that environment and inspect relevant logs; deploy success alone is not behavior proof | |
Docs only (.md, AGENTS.md, SKILL.md) | Nothing to run | Say "docs-only, no runtime check applies" — don't invent a verification step |
For any user-visible change, put the proof in the reply: a screenshot of the surface you just drove, or the actual query result / log line for backend work. "Show me screenshots" is a standing expectation, not a special request.
pnpm test:fast excludes .db.test.ts / .integration.* / .e2e.* /
.live.* / .perf.* suites. If your change touches one of those, name and
run that specific file — test:fast passing does not cover it.
When inspecting production runs, query interactive and scheduled/background
work as separate populations before summarizing reliability. Report both
id NOT LIKE 'job-%' and id LIKE 'job-%' (or the repo's current equivalent),
including app, run count, completed count, failure count, and top terminal
reasons for each slice. A healthy interactive sample does not prove scheduled
jobs work.
State it plainly and name what would close the gap: "I could not run this —
verifying it needs <command> or a browser check of <page>." Never write
"should be fixed" or "this resolves it" without having actually run the check
above.
f07726b
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.