Server chat state with withPersistence from @tanstack/ai-persistence. Authoritative transcript, run lifecycle, durable interrupts/approvals, chatParamsFromRequest, reconstructChat, snapshotStreaming. Use when the server owns history, multi-device, or durable tool approvals. NOT client localStorage (see ai-core/client-persistence in @tanstack/ai) and NOT stream reconnect alone.
Builds on ai-persistence. Package:
@tanstack/ai-persistence.
withPersistence(persistence) is a ChatMiddleware that writes chat state
to a backend: messages, runs, interrupts (optional metadata). It does not
mutate the chunk stream and does not replace delivery durability.
import {
chat,
chatParamsFromRequest,
toServerSentEventsResponse,
} from '@tanstack/ai'
import { openaiText } from '@tanstack/ai-openai'
import { withPersistence } from '@tanstack/ai-persistence'
// Your adapter — see ai-persistence/stores.
import { persistence } from './persistence'
export async function POST(request: Request) {
const params = await chatParamsFromRequest(request)
const stream = chat({
adapter: openaiText('gpt-5.5'),
messages: params.messages,
threadId: params.threadId,
runId: params.runId,
...(params.resume ? { resume: params.resume } : {}),
middleware: [withPersistence(persistence)],
})
return toServerSentEventsResponse(stream)
}Always pass threadId and runId from the client (via
chatParamsFromRequest / body helpers). Forward resume when the client
resolves pending interrupts.
For dev and tests, memoryPersistence() from @tanstack/ai-persistence is a
drop-in backend that implements all four stores in process.
| Store | Role | Required? |
|---|---|---|
messages | Full model-message transcript load/save | Yes for withPersistence |
runs | Run status, timing, usage, errors | Optional; needed for interrupt durability |
interrupts | Pending/resolved tool approvals & waits | Optional; requires runs |
metadata | App-owned namespaced key/value | Optional |
Named shapes: ChatTranscriptPersistence (floor), ChatPersistence (all four).
Annotate your factory with one of these, not with bare AIPersistence —
the unparameterized type is the all-optional bag, and withPersistence rejects
it because stores.messages is possibly undefined.
withPersistence merges incoming messages into the stored thread by id.
messages: load the stored thread and continue.messages: merge by id. The last incoming id that already
exists in stored is a cutoff. Stored messages after it are dropped. If no
incoming id is in stored, every stored message stays. Same id: incoming
wins. New ids and messages with no id are appended.saveThread replaces the thread with that merged list. Merge is middleware,
not the store.| Moment | Writes | Best-effort? |
|---|---|---|
onStart | Pending turn snapshot (user + history) | Yes — failure does not abort |
| Interrupt boundary | New interrupts, run → interrupted, message snapshot | No |
onFinish | Canonical transcript first, then run → completed, commit resumes | No |
| Stream (optional) | Throttled partial assistant text | Yes if snapshotStreaming: true |
onError | Run → failed | Resumes stay pending |
onAbort | Run → aborted — but only sometimes (see below) | Resumes stay pending |
The canonical transcript already contains the completed terminal assistant messages. Native-combined output keeps the structured result on its terminal assistant message. Separate finalization and event-sourced harness output can preserve plain-text and structured-output assistant messages separately when those messages use different ids.
import { withPersistence } from '@tanstack/ai-persistence'
import { persistence } from './persistence'
withPersistence(persistence, {
snapshotStreaming: true,
snapshotIntervalMs: 1000, // default
})onAbort writes conditionally, not alwaysA user pressing Stop and a user closing the tab produce the identical
connection close, so onAbort can never infer intent from the abort alone.
It writes:
'aborted' (terminal, with finishedAt) when the abort is an explicit
cancel — info.cancelRequested === true, or a durable cancel request found
via wasCancelRequested(runs, runId) (both from @tanstack/ai; paired with
requestRunCancel/RUN_CANCEL_REASON) — or when the run is not
detachable at all (no sandbox/journal behind it, so there is nothing to
reattach to).@tanstack/ai-sandbox, has provided
DetachableRunCapability from @tanstack/ai). The record deliberately
stays 'running' — the agent keeps running and a later attach can take it
over. (The detaching middleware, not withPersistence, is what stamps
detachedSince.)Chat's onAbort and generation's onAbort (withGenerationPersistence) are
asymmetric on purpose: a generation job has no journal and no agent loop
to reattach to, so its onAbort always writes 'aborted' unconditionally.
Do not "fix" that asymmetry by making generation conditional, or chat
unconditional — both are correct for what they wrap.
Never build a client, or a persistence backend, that assumes a disconnect
always finalizes the run — for a detachable run it usually does not, and
inventing a finishedAt for a still-'running' record breaks takeover.
Use isTerminalRunStatus(status) (from @tanstack/ai-persistence) to test
whether a status is finished, rather than re-listing
'completed' | 'failed' | 'aborted' by hand.
Streaming snapshots default off (finish is authoritative). Enable only when partial-output durability is worth extra writes.
Resumes accepted in onConfig commit only at a success boundary (interrupt or
finish). A failed run leaves interrupts pending so the same resume batch can
retry.
resume batch or onConfig
throws.resumeToolState and clears
config.resume so the engine does not double-reconstruct from client
history (server owns transcript).reconstructChat)Server-authoritative clients load history by threadId (often GET):
import { reconstructChat } from '@tanstack/ai-persistence'
import { persistence } from './persistence'
import { sessionUserId, userOwnsThread } from './auth'
export async function GET(request: Request) {
return reconstructChat(persistence, request, {
// Multi-user: required in production
authorize: async (threadId, req) => {
const userId = await sessionUserId(req)
return userOwnsThread(userId, threadId)
},
})
}Returns { messages, activeRun, interrupts, page? }:
messages: UI messages for this windowactiveRun: { runId } if a run is still generating (runs.findActiveRun)interrupts: pending human-in-the-loop state for re-promptpage: { truncated, cursor } when the GET has a valid limitPaging is opt-in. No limit returns the full transcript and can omit page.
reconstructChat reads limit and before from the query. activeRun and
interrupts are not paged.
Without authorize, anyone who guesses ?threadId= gets the transcript.
withGenerationPersistence(persistence) tracks run records for non-chat
activities (image, audio, TTS, video, transcription). Do not fake
threadId = requestId on chat run stores — use the generation helper.
saveThreadMerge by id is withPersistence. saveThread must replace the merged list it
receives.
threadId / runIdPersistence keys and resume need stable ids. Use chatParamsFromRequest.
runsinterrupts requires runs; withPersistence throws otherwise.
AIPersistenceAIPersistence defaults to the sparse all-optional bag, so withPersistence
and reconstructChat reject the value. Return ChatPersistence (or
ChatTranscriptPersistence) instead.
withPersistence to reconnect a dropped streamThat is delivery durability (resumable streams), not state persistence.
@tanstack/ai) — browser half7fb4a5f
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.