Authors and calibrates Instance AI evaluations that build standalone n8n Agents through Agent Builder. Use when a change under packages/cli/src/modules/agents affects build-agent routing, Agent setup, model or credential selection, tools, MCP servers, integrations, skills, tasks, testing, or user-facing build responses. Requires LangTracer access before authoring so each finished case can be published.
Use the shared Instance AI eval harness. Agent cases use an Agent-specific authoring directory, dataset, and LangTracer suite.
Run this check before sourcing, drafting, or writing an eval. Run it from
packages/@n8n/instance-ai:
pnpm exec dotenvx run -f ../../../.env.eval -- \
sh -c 'test -n "${LANGTRACER_URL:-}" && test -n "${LANGTRACER_API_KEY:-}"'If the check fails, stop before creating an eval file. Ask the user to:
Generate a key on the LangTracer API page.
Add these variables to the repository root .env.eval file:
LANGTRACER_URL=https://lang-tracer.n8n-maintenance.workers.dev
LANGTRACER_API_KEY=<generated-key>Confirm when the environment is ready.
Do not ask the user to paste the key into chat. Do not print or inspect its value. Rerun the check after the user confirms. Continue only when it passes.
packages/@n8n/instance-ai/evaluations/data/agents/<slug>.json."datasets": ["agents"].agents.The disk runner loads both data/agents/ and data/workflows/. A misplaced
Agent case can therefore pass locally. That does not make the location correct.
Write the smallest user request that exercises the changed behavior.
processExpectations for the Instance AI conversation and final response.outcomeExpectations for the created Agent artifact and its configuration.executionScenarios only when the built Agent must run to prove the behavior.The harness captures the Agent configuration and authored skills. It supplies them to the expectation judge. A scenario-less Agent case is valid when process or outcome expectations can prove the behavior.
Use the substitution test for every expectation. A correct alternative build must pass. A build that misses the requested behavior must fail.
Agent Builder runs as a delegated sub-agent architecture today. That design can change. Do not write an expectation that depends on it. Follow "Keep expectations free of internal mechanics" in create-instance-ai-eval.
Do not write:
build-agent" or "calls build-agent once."call_agent to verify the Agent."Write what a user or a reviewer can observe:
| ❌ mechanical | ✅ intent |
|---|---|
| "The sub-agent tests the Agent before it reports" | "The final response says the Agent was tested only if a test ran, and it reports the result of that test honestly" |
"build-agent is called with the Slack credential" | "The Agent uses the Slack credential the user has" |
| "The sub-agent adds a tool for the lookup" | "The Agent can look up the order status when a user asks for it" |
"The builder sets model after the catalog lookup" | "The Agent uses a model that the connected credential supports" |
Put each one where it belongs:
outcomeExpectations.processExpectations.The calibration steps below ask you to confirm that build-agent ran. That is
a check for you, to be sure the case reached Agent Builder. Do not add it to the
case.
For multi-turn, seeded, or capability-gap cases, follow the case-shape and calibration rules in create-instance-ai-eval. This skill overrides its workflow directory and suite guidance for Agent cases. The sections below add the Agent-specific parts. Do not skip the general rules they point to.
Follow "Ask for the target suite first" in create-instance-ai-eval. Do it after the LangTracer preflight and before you source or draft anything. If the request already names a suite, state it and do not ask. Otherwise ask, in autonomous mode too, in the same message as the autonomy question.
For Agent cases, recommend
Instance AI capabilities — agents
(slug agents). Offer baseline as well: the Instance AI (INS) team monitors
it and it runs nightly. Get the full list from list_suites. Use the suite the
driver picks in every push command, in place of agents in the examples below.
Keep the agents dataset and the data/agents/ directory. They do not depend
on the suite.
Follow "Set the autonomy level first" in create-instance-ai-eval. Ask the driver for autonomous or checkpoint mode before you source or draft anything. The four gates are the same: selection, shape and expectations, calibration, push.
The best Agent cases come from a real conversation. Discover the conversation with LangTracer. Author a synthetic case from what you learn.
list_conversations with usedAgentBuilder: true. Do not
use source: "agent-builder". That pool is a different product, has
unscrubbed payloads, and has no Agent snapshots.list_conversation_agent_snapshots.
target-resolved is the state that a turn opened on. config-updated is the
state after the builder changed it.get_conversation_agent_snapshot. Seed from the
target-resolved row of the turn under test.Two snapshot patterns to look for:
target-resolved hash that differs from the previous config-updated hash.
The user edited the Agent in the UI between turns. This is the best seed for a
repair or update case.config-updated. The user added one rule per
turn. Use it for a "new rule does not erase old rules" case.Credentials arrive as [redacted]. Instructions and skill bodies do not. Scrub
the prose before it enters a case. See the agents section in
case-shapes.md (the agents seed section).
| Shape | Question it answers | How to write it |
|---|---|---|
| Create | Does the builder make the right Agent from a fresh request? | One user turn. No seed. Assert on the Agent artifact. |
| Update | Does the builder change only what the user asked for? | seed.agents with the Agent before the change. One live turn with the change. Assert on the new behavior and that the untouched parts survive. |
| Repair | Does the builder find the cause of a failure and fix it? | Seed the broken Agent. The live turn reports the failure. Assert on the cause found and the fix applied. |
| Incremental rules | Do earlier rules survive when a new rule lands? | Seed the Agent after several rules. Add one conflicting or overlapping rule. Assert that the old rules are still in force. |
Rules that apply to all four:
conversation. Put earlier turns in seed.messages.A pass can mean the builder did the right thing. It can also mean the situation never happened. Check before you trust a green.
Follow "A red is signal" in the general skill. For Agent cases, add this:
description with
Capability-gap finding:.Agent. Propose the ticket. Do not create it unless the driver says go.
Follow "Capability gap → propose a Linear ticket" in the general skill. Also
check whether the gap already has a ticket.
packages/cli/src/modules/agents/builder/agents-builder-prompts.ts,
packages/cli/src/modules/agents/builder/prompts/, and
packages/cli/src/modules/agents/builder/skills/. Also read the Instance AI
orchestrator prompt in packages/@n8n/instance-ai/src/agent/system-prompt.ts.
If none of them covers the failing situation, add that to the ticket as a
proposed prompt or skill change.--iterations 5.When the failure is in the Agent's runtime (for example a channel that cannot
report an error), the builder cannot see it. The right fix is often a builder
instruction to say so and stop guessing. Write that as a processExpectations
item on the final response.
Follow "Share links, never bare ids" and "Link the pushed case to its source" in the general skill. For a sourced Agent case:
https://lang-tracer.n8n-maintenance.workers.dev/test-cases/<id> form.update_test_case with sourceThreadId, expectedBehavior, and
failurePattern.add_case_tags.--set-kind capability_gap into a
suite of that kind. If no such Agent suite exists, tell the driver. Do not put
it in a regression suite.Start with this shape:
{
"description": "The Agent Builder behavior this case guards.",
"conversation": [
{ "role": "user", "text": "Build me an Agent that ..." }
],
"complexity": "simple",
"tags": ["agent", "agent-build", "<capability>"],
"credentials": [{ "type": "<credentialType>", "name": "<display name>" }],
"processExpectations": [
"The final response ..."
],
"outcomeExpectations": [
"A standalone Agent is the deliverable; any workflow created exists only as a tool the Agent calls.",
"The Agent ..."
],
"datasets": ["agents"]
}Keep the prompt in the user's voice. Do not tell Instance AI which internal tools or configuration fields to use unless that choice is the behavior under test.
Read local-setup.md when the machine does not already have an eval instance and environment file.
From packages/@n8n/instance-ai:
pnpm exec tsx -e "import { loadAgentEvalTestCasesWithFiles } from './evaluations/data/agents/index.ts'; const matches = loadAgentEvalTestCasesWithFiles('<slug>'); if (matches.length !== 1) throw new Error('Expected exactly one Agent eval case, found ' + matches.length); console.log(matches[0].fileSlug)"
pnpm eval:instance-ai \
--base-url http://localhost:5680 \
--filter <slug> \
--tier agents \
--concurrency 1 \
--keep-workflows \
--verboseUse eval:instance-ai for a new disk case. eval:agents reads the published
LangTracer suite and is for running cases that are already there.
Inspect the transcript, the rendered Agent artifact, and each judge reason. Do not accept a green result when a conditional expectation never occurred. Do not weaken an expectation to hide a real Agent Builder defect.
The Instance AI PR gate checks the files changed by the PR. A change under
packages/cli/src/modules/agents/ selects the Instance AI capabilities — agents
suite through its agents slug. It also selects the agents dataset and uses
an absolute pass gate. The run uses a suite-scoped LangSmith cohort. It does not
write to the workflow dataset or compare against the workflow baseline. Other
Instance AI changes select the baseline suite and its pr dataset.
The gate runs when a PR opens, reopens, or becomes ready for review. It does not run for each new push. Use the PR gate's manual dispatch after a later push.
Declared credentials are real n8n credential records with placeholder data. The eval thread limits the builder to those credential IDs.
Agent Builder model catalog requests return deterministic fake models during an eval. They do not decrypt the placeholder model credential or call its provider. Production model catalog requests remain live.
This mock covers catalog lookup only. A builder call_agent action and an Agent
executionScenario run the target Agent model. They need a working provider
credential such as EVAL_OPENAI_API_KEY. A build-only case does not need one.
--concurrency 1.build-agent.--iterations 5 before adding a case to a gating tier.<suite> below).pnpm exec dotenvx run -f ../../../.env.eval -- \
pnpm eval:langtracer-push --suite <suite> --dry-run --changed
pnpm exec dotenvx run -f ../../../.env.eval -- \
pnpm eval:langtracer-push --suite <suite> --changedThe push needs LANGTRACER_URL and LANGTRACER_API_KEY. Generate a key on the
LangTracer API page.
Report the case and suite as clickable LangTracer links. Delete the local JSON
after a successful push.
data/agents/ and uses the agents dataset.agents suite in CI. A case pushed to
another suite does not run in that lane.0e1c754
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.