Write or extend an E2E journey spec in evals/specs that proves a PR's change to a human reviewer. Use when a PR changes user-visible behaviour, when a reviewer asks "show me", or when the coverage decision says a journey is missing.
Every evals/specs/**/*.e2e.test.ts a PR adds or changes runs in CI on the PR
head and is published as that PR's proof (private report + one PR comment).
Write the spec for the person who will read that report in thirty seconds,
not for the machine that runs it.
| Moment | Skill | Output |
|---|---|---|
| Decide whether a journey is needed and design it | write-a-spec (this) | a spec whose titles, steps, and evidence read as the report |
| Run it and read your own record | run-tests | evals/results/test-runs/<latest>/ with the captions in order |
| Let a reviewer click through the change | preview-my-work | a sandbox link in the report's Show sandbox panel |
| Describe it on the PR | open-a-pr | an ## Evidence line naming the spec and its before → after |
| A proof or check is red | diagnose-a-red-run | classification before any code change |
| Extra screenshots or a video outside the spec | record-a-demo, upload-photo | PR attachments (these are not proof) |
CI picks the lane from the file name and tags alone. Pick the lane first: it decides the world you can build.
| If the spec… | Name / tag | CI lane |
|---|---|---|
| runs in a browser or local world with no secrets (default) | evals/specs/<journey>.e2e.test.ts | PR change proof, unprotected |
| must boot the packaged desktop binary | evals/specs/packaged-<journey>.e2e.test.ts | packaged smoke runner |
| needs a reviewer to reopen the running browser at a screenshot | add { tags: ["checkpoints"] } to the test(...) options | protected checkpoint lane on Freestyle |
| needs Windows and a published installer | the existing windows-published-preview spec only | Daytona Windows, protected |
Tag checkpoints only when a still image cannot show what the reviewer has to
check: an open stream, a populated workspace to explore, a state that costs
minutes to reach. The tag moves the spec to a protected lane that uses a
secret, where fork PRs cannot run. It saves the end state of each tagged test,
plus every user.checkpoint("caption") and every step(..., { checkpoint: true }).
Plain user.screenshot() never saves one. Say why in a comment beside the
tag. Checkpoints cost VM snapshots, so capture the one or two moments a
reviewer would open, not every step. Run it locally with
pnpm evals:e2e <slug> --local --checkpoints (needs FREESTYLE_API_KEY).
../../apps|packages|ee, read source files, or spawn another
runner. That is a unit test in disguise; the boundary ratchet rejects it.No UI does not mean no spec. A server, proxy, or protocol bug gets a
browser-less world (see evals/specs/session-title-recovery.test.ts); the
steps still read as one person's before → after.
A proof answers: who can now do what they could not before, and who
still cannot. Structure every spec as that story. Each beat is one step()
with a claim-sized name; each visual beat ends in user.screenshot().
persona → the test title names a person: "an owner", "a teammate", "a member without access"
before → what that person sees or cannot do today (features only)
action → the person does the thing, through the UI, as they would
after → what they now see; screenshot at that moment
boundary → who else is affected, and the negative half: who is notThe review app is generated from the spec. It shows exactly four strings, and each comes from one place in your code; write those strings for the reviewer:
| Reviewer sees | Comes from | Rule |
|---|---|---|
| Section heading | the test("…") title | names the persona and what they can now do |
| Caption under a screenshot | the step("…") the user.screenshot() ran inside | the old state starts before:, the new state after:; otherwise a plain claim |
Caption + judgment on a looks() image | the first expectation in user.looks([...]) | judged later; pending until then, so CI proof stays Incomplete |
| Assertion line | recordAssertionEvidence(claim, evidence, ok) | claim is the caption, evidence the text under it |
Each test is one section, marked with a status icon. Its checks collapse to
"N of M checks passed" and open only when one fails. Its screenshots form a
numbered gallery in capture order, so the captions have to tell the story on
their own. A section with no recordAssertionEvidence shows as Incomplete.
A screenshot taken outside any step() is captioned " artifact N",
which tells the reviewer nothing. No verbs like "assert", no selectors, no
internal names anywhere in those strings.
Example title and steps:
test("an owner enables Code Mode and a teammate turns one chat into a shared Workflow", async ({ user, probe, step }) => {
await step("before: the teammate's agent has no script tool", …); // screenshot
await step("the owner enables Code Mode with one switch", …); // screenshot
await step("after: the teammate's request runs as one script", …); // screenshot
await step("the result is saved as a Workflow the team can open", …); // screenshot
await step("a member outside the team cannot see it", …); // screenshot
});Assertions still live inside each step; they just hang off user-visible moments. A spec whose assertions are RPC responses collected in the world and compared in one block is a mechanism check, not a proof. It is still allowed; it just tells the reviewer nothing.
The spec, the report, and the PR body tell one story in the same words. Write
the report you want to read first, then make the code produce those lines.
A bug fix with no screen still gets this shape; the "screenshot" is one
recordAssertionEvidence line per step whose evidence is the observed fact.
Report the reviewer reads (every line comes from the spec):
▶ a member's conversation stays readable while OpenWork is still checking who owns it
✔ given an engine where the ownership check is slow and the messages read is fast
GET /session/:id answers in 150 ms; GET /session/:id/message in 2 ms
✔ when the member opens the conversation and memory is reclaimed mid-check
three reads sent as the app sends them; gc() ran three times
✔ then the runtime really reclaimed the engine's response
WeakRef → undefined (witness; without it a pass means "nothing happened")
✔ after: all three reads return the conversation, none say internal_error
200 / 200 / 200; body {"id":"ses_x", …}Spec that produces it:
test("a member's conversation stays readable while OpenWork is still checking who owns it", async ({ world, step, evidence }) => {
await step("given an engine where the ownership check is slow and the messages read is fast", async () => {
evidence.recordAssertionEvidence("engine timings", `GET /session/:id ${world.ownershipMs} ms; GET /session/:id/message ${world.messagesMs} ms`, true);
});
await step("when the member opens the conversation and memory is reclaimed mid-check", async () => { /* act; gc(); */ });
await step("then the runtime really reclaimed the engine's response", async () => {
const collected = world.engineResponse.deref() === undefined;
evidence.recordAssertionEvidence("engine response reclaimed", collected ? "WeakRef → undefined" : "WeakRef still alive: window not exercised", collected);
expect(collected).toBe(true);
});
await step("after: all three reads return the conversation, none say internal_error", async () => {
evidence.recordAssertionEvidence("reads", `${statuses.join(" / ")}; body ${body.slice(0, 40)}`, statuses.every((s) => s === 200));
expect(bodies.join()).not.toContain("internal_error");
});
});PR body line that names it (open-a-pr):
## Evidence
`evals/specs/conversation-stays-readable.e2e.test.ts` — before: the read fails
with "Response body object should not be disturbed or locked"; after: all
three reads return the conversation and the reclaim witness is true.Rules that make this parseable:
given / when / then / after: (or before: / after: with
screenshots). One clause each, the person's words, no function names.dev and green on the head; never narrate it.Rules of thumb:
user.* act or user.see. probe.* observes;
it does not carry a step alone.recordAssertionEvidence
with the command or response excerpt as the evidence text.run-tests → Match the runtime).seed.appWeb (headless Chrome, real app). Use seed.desktop only for
a native capability a browser cannot show, and say why in nativeReason.seed.evalIn / probe.eval in new specs. If you need one, comment why.Import spec (and expect) from @openwork/testkit; bind the world:
const test = spec.world(myWorld, {
resources: { surfaces: ["appWeb"], services: ["den"] },
});| Channel | Use it for |
|---|---|
seed | Arrange the world: Den, orgs, members, workspaces, sessions, mocks, faults. Only seed writes state. All seed.* goes in the world, before the first act. |
user | Act as the person: click, type, press, reload, see, notSee, looks, screenshot. Trusted CDP input; no JS evaluation. |
agent | Drive the product's automation rail (window.__openworkControl): sends, session actions. |
probe | Observe without changing state: text, hash, storage, api (GET), dom, eventually. |
step | Name a claim. Nests. A failed step is recorded and rethrown; later steps show not-reached. |
Well-known targets: "composer" for the editor; otherwise role, label, text,
placeholder, or test id. Bound every wait; declare external requirements in
needs() so a missing dependency skips loudly.
mcpMock() witnesses under evals/packages/labs/src/; never call real
providers. Witnesses are deterministic, identity-scoped, and queryable.Evidence is ambient: user.screenshot() records an artifact, user.looks()
records a visual validation (judged later; pending until then),
recordAssertionEvidence() records a witness assertion. Never create or
pass recorder handles.
Run it once before pushing and read your own record:
pnpm evals:e2e <slug> --local
ls evals/results/test-runs/<latest>/ # index.html, test-run.json, NN-*.pngIf the screenshots would not convince you, they will not convince the reviewer.
Read the captions in index.html top to bottom: they should tell the
before → after story on their own.
In CI the spec runs on PR change proof, one job per spec; the trusted
publisher aggregates every changed spec's records into one report. Failed,
skipped, and cancelled runs stay visible as such; nothing substitutes for them.
417244c
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.