Reproduce and measure a Mastra Code runtime bug against a real model by driving the built TUI headlessly in tmux while every layer (network, provider stream, run engine, TUI) appends timestamped JSONL. Use when a bug depends on real provider timing, streaming order, or event flow that mocked tests and fixtures may not reproduce — e.g. wrong token rates, flicker, stuck status, ordering races, or "only happens sometimes with a real model".
77
96%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
Run the real TUI against a real model, have an agent type the prompts, and log timestamps at each layer. Comparing the layers step by step shows where timing or ordering goes wrong: at the provider, in the network wrapper, in stream conversion, in the run engine, or in the UI reducer.
Pair this with debugging-difficult-bugs: instrument, reproduce, read the log before fixing, and clean up afterwards. Afterwards, turn the finding into deterministic checked-in coverage (unit tests plus a TUI E2E scenario under mastracode/tui/e2e/tui/). A live run is evidence, not a regression test.
The TUI uses the user's configured model and stored credentials (for example a Claude Max OAuth login), so every prompt spends their quota. Use short prompts. Tell the user which model you ran and roughly how many sessions.
Append one JSON line per event to <cwd>/debug-token-rate.jsonl (or a similarly named file) with appendFileSync. Rules:
build marker (for example 'my-bug-diagnostics-v1'). That lets you confirm the build you ran contains the instrumentation and find every leftover when cleaning up.try {} catch {} so logging can never break the run.Useful hook points, from outermost to innermost:
| Layer | Where | What to log |
|---|---|---|
| Network | The provider's fetch wrapper, e.g. buildAnthropicOAuthFetch in mastracode/sdk/src/providers/claude-max.ts | request start; headers arrival (status, content-encoding); each body chunk via a pass-through TransformStream (bytes, SSE event: names). Re-wrap the body with new Response(body, { status, statusText, headers }). |
| Provider stream | convertFullStreamChunkToMastra in packages/core/src/stream/aisdk/v5/transform.ts | raw chunk type and delta length as the AI SDK hands it over |
| Run engine | processStreamChunk in packages/core/src/agent-controller/session-run-engine.ts | chunk type, getChunkProducedAt(chunk), current message id, step-start startedAt, usage. Skip per-delta chunks here if the provider layer already logs them. |
| TUI | mastracode/tui/src/tui/event-dispatch.ts (web mirror: mastracode/factory-ui/src/ui/domains/chat/services/runtime.ts) | event type, ids, relevant state before and after, and any computed value (e.g. usage.computed with tokens, window, instantaneous and displayed rate) |
For other providers, find the equivalent fetch option or wrapper in mastracode/sdk/src/providers/.
pnpm install --frozen-lockfile # fresh worktrees only
pnpm build:mastracode # ~50s once deps are built
rg -l "my-bug-diagnostics-v1" mastracode/tui/dist packages/core/dist mastracode/sdk/distIf rg finds nothing in a dist, the run will not produce that layer's records.
Start the TUI in a throwaway git repo so the agent's tool calls default to scratch files. Setting the working directory is not a sandbox: do not grant the session extra allowed paths, and keep prompts pointed at the scratch repo.
mkdir -p /tmp/mc-repro && git -C /tmp/mc-repro init -q
tmux new-session -d -s mcrepro -x 200 -y 50 -c /tmp/mc-repro \
"node $PWD/mastracode/tui/dist/cli.js"
sleep 14 # startupSend each prompt, then send Enter as a separate key press, then wait and read the screen:
tmux send-keys -t mcrepro "/new"; sleep 1; tmux send-keys -t mcrepro Enter; sleep 3
tmux send-keys -t mcrepro "Please do these one at a time as separate tool calls (not in parallel): write 3 files named a1.md through a3.md, each with a ~100 word paragraph about a different tree, writing a sentence of explanation before each tool call. Then run ls -la with execute_command and summarize in three sentences."
sleep 1; tmux send-keys -t mcrepro Enter
sleep 50
tmux capture-pane -t mcrepro -p | grep -v '^\s*$' | tail -3 # status line is lastTips:
/new, because intermittent behavior needs a sample. Record what the screen showed each time so you can check the log against what the user would have seen.dist.Group the log records by step (for example, split at each usage.computed or step-finish record) with a short python3 script. For each step, compare durations across layers: request → headers → first content → last content at the network layer, and the same span at the provider, engine, and UI layers.
accept-encoding: identity), rebuild, and rerun.Report the evidence as a small table (normal steps vs. anomalous steps), along with the model used, the number of runs and steps, and what the data cannot prove.
tmux kill-session -t mcrepro
rg -n "my-bug-diagnostics" mastracode packages --glob '!**/dist/**' # must print nothing
git diff # read it: only intended changes remain
rm -f debug-token-rate.jsonl /tmp/mc-repro/debug-token-rate.jsonlAlso revert any experiment toggles (like changed request headers) and confirm with rg before saying they are gone. Rebuild so dist no longer contains the instrumentation.
61079f5
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.