Review Agent-Written Code Against Your Team's StandardsLearn More
Book a Demo
CareersDocs
Log inBook a Demo

ARTICLE

AI Agent Evaluation Needs Executable Specs

Why AI coding agents need executable specs, scoped verification, and safe browser sandboxes to turn product intent into trustworthy software checks.

Shachar Azriel

Shachar Azriel

·17 Sept 2026·12 min read

Code review is not slow because engineering teams forgot how to review code. It is slow because the amount of code we can generate has outgrown the amount of trust we have in the output.

I see this every week. I speak with engineering teams ranging from five-person startups to organizations with thousands of developers. The pattern is consistent: AI coding agents are helping teams produce more pull requests, but those pull requests are spending more time waiting for review. Velocity was supposed to improve. In many places, it is simply moving the bottleneck.

At AI Native DevCon London, I talked about why this happens and what we can do about it. The question behind the talk was simple: in 2026, why are AI code reviews still a bottleneck in engineering teams?

The short answer is trust. Coding agents are getting better at generating code, but they are not yet consistently focused on verifying whether the feature matches the original product intent. If we want agentic coding to scale inside real teams, we need a verification layer that sits between the specification and the shipped behavior.

Use This Talk As Agent Context

If you want to put the ideas from my AI Native DevCon talk to work, Tessl has turned the talk into a skill your agent can use as context. You can also watch the full recording.

DevCon NYC
Register to get the early birds discount

AI Agent Evaluation Has To Start With Intent

AI agent evaluation is the practice of checking whether an agent has achieved the intended outcome, not merely whether it produced plausible code. For coding agents, that means comparing implemented behavior against the feature specification, product constraints, visual requirements, and important user flows.

That distinction matters because agents can look successful while still being wrong.

In one recent example, I asked an agent to add an integration to an onboarding flow. I also attached a previous recording that showed a specific UI bug: a button was overlapping another element. My instruction was explicit. Do not repeat this issue.

When the developer sent me the preview environment, the continue button was exactly where I had asked it not to be.

This was not a complex distributed systems problem. It was a straightforward frontend task with a visible failure case. If an agent can miss that, we should expect far more difficult failures when it works across permissions, payments, onboarding, integrations, and business logic.

The deeper problem is that our workflows often ask humans to detect those mistakes after the agent is done. That leaves product managers, engineering managers, and reviewers as the verification system.

Executable Specs Turn Product Intent Into Checks

Executable specs are specifications that an agent can convert into concrete verification work. They are not only documents for humans. They become the input for an agent that can extract requirements, inspect a preview environment, compare behavior against intent, and return a verdict.

Most good engineering teams already have the raw material. It lives in Jira, Linear, Monday, Notion, GitHub issues, Figma files, product docs, support notes, and pull request descriptions. The missing piece is a system that can test whether the implementation respects that intent.

The first version of our spec reviewer tried to do this with one agent. It had access to the feature specification, the designs, and the deployed preview environment. The goal was simple: extract all requirements and check whether they were implemented.

It quickly hit a limit. One agent trying to understand the spec, plan the work, navigate the application, inspect the UI, and produce a useful report can run out of context or lose consistency. The solution was to split the work.

One agent became the planner. Its job was to read the specification and identify requirements, sub-requirements, edge cases, and potential failure modes. Other agents became verifiers. Their job was narrower: inspect the relevant part of the application and check one requirement at a time.

That changed the quality of the output. Instead of one general agent trying to review everything, we could run many focused agents in parallel. Each one returned a verdict. An orchestrator then collected the results into a report a team could actually use.

Specs Plus Code Are The Verification Layer

The next lesson was that specs alone are not enough.

If you give an agent only the specification and design, it can understand the ideal outcome, but it may also invent requirements nobody asked for. That creates noise, and noise is dangerous in review workflows. Once reviewers stop trusting the report, the system becomes another thing to ignore.

The better approach is to ground the agent in both the specification and the code. Specs are intent. Code is reality. Together, they are a gold mine.

One surprising finding was that reviewing against the base branch can be more useful than reviewing only the diff. A diff shows the implementation path the developer or agent chose. That can bias the reviewer. The base branch helps the verification system stay closer to the product question: given what already exists, what should this feature change, and does the preview now behave that way?

Scope matters too. If the change is a frontend feature, the agent should not generate irrelevant backend findings just because it can inspect backend files. Verification needs boundaries. It should know what kind of feature it is reviewing, which surfaces matter, and which failures would actually block the team.

This is where executable specs become practical rather than theoretical. They let teams move from "please review this PR" to "check whether this implementation satisfies the stated requirements and critical flows."

Verification Has To Run Where Risk Is Contained

There is another issue that teams should not ignore: security.

To verify preview environments, an agent may need to click unknown links many times each month. In a normal security conversation, we would describe that as phishing behavior. Now imagine letting that agent run near your source code, customer data, credentials, cloud storage, and model keys.

That is not a small operational concern. It is a real architectural risk.

Our conclusion was that we should not build dangerous security infrastructure casually inside the product. For high-risk parts of the SDLC, it is often better to use a proven third-party sandbox than to improvise. In our case, each requirement could run inside an isolated browser session, return a verdict, and avoid exposing the rest of the environment.

This pattern also applies beyond validating new requirements. Customers quickly asked about regression for critical flows. If a subscription flow breaks, a company can lose customers immediately. I gave the example of watching a customer try to subscribe with 100 seats, fail, and leave the product. If onboarding breaks, the user may never reach activation. If a dashboard starts showing inconsistent data, the user may lose trust in the whole system.

That is why our agents were not only checking whether a new feature matched its ticket. They were also navigating dashboards, trying integrations, walking through subscription flows, and testing login and onboarding paths. Those flows are not every possible thing the product can do. They are the flows where breakage is costly enough to justify repeated verification on every pull request.

Not every part of the product deserves expensive agentic verification, but the most important one to five percent often does.

That is the practical version of AI agent evaluation. It is not "let an agent explore the whole application forever." It is "point agents at the flows, requirements, and risks that matter most."

The Opportunity Is In The Gaps

The strongest lesson from this work is that context engineering is still hard. Even in 2026, complex agentic workflows do not work out of the box just because the model is capable. They need planning, delegation, scoped context, reliable verification, and safe execution environments.

The good news is that most teams do not need to invent a new source of truth. They already have specs. They already have code. They already have preview environments. The opportunity is to connect those assets into a verification loop that helps teams trust what agents produce.

That is also why I do not think the answer is simply to generate hundreds of tests for imaginary scenarios. Tests are useful, but the scenarios that matter most are often already described in the work the team planned: the product spec, the design, the ticket, and the critical flows customers actually use. Executable specs start from that intent and then check the running product against it.

For startups and tooling teams, this is also where the market gets interesting. Big coding agents are improving quickly, but they leave gaps. The best opportunities are often where teams need a specialized layer: verification, workflow integration, safety, reporting, or product-specific evaluation.

That is how I think about executable specs. They are not just better documentation. They are a way to turn product intent into something agents can act on, test, and report back to humans.

If you want the full walkthrough, including the architecture decisions, failures, security tradeoffs, and Q&A, you can watch the recording.

COPY & SHARE

Shachar Azriel

Shachar Azriel

For the past decade, I’ve helped startups and mid-sized tech companies scale teams, establish systems, and launch products that stick. Today, I’m VP of Product at Baz, where we’re on a mission to reinvent code review with AI: making it faster, smarter, and (believe it or not) more fun for developers. We’re working at the bleeding edge of technology, facing unique challenges that many other product and development teams are only beginning to encounter. That’s why I regularly share real stories from building AI-powered features in the wild, and from the journey of building an AI-driven company itself. Beyond the product, I love connecting with people. I co-founded the AI-Dev community in Israel, dedicated to accelerating innovation in AI coding and product development.

READING

·

0%

IN THIS POST

Use This Talk As Agent ContextAI Agent Evaluation Has To Start With IntentExecutable Specs Turn Product Intent Into ChecksSpecs Plus Code Are The Verification LayerVerification Has To Run Where Risk Is ContainedThe Opportunity Is In The Gaps

COPY & SHARE

Shachar Azriel

Shachar Azriel

For the past decade, I’ve helped startups and mid-sized tech companies scale teams, establish systems, and launch products that stick. Today, I’m VP of Product at Baz, where we’re on a mission to reinvent code review with AI: making it faster, smarter, and (believe it or not) more fun for developers. We’re working at the bleeding edge of technology, facing unique challenges that many other product and development teams are only beginning to encounter. That’s why I regularly share real stories from building AI-powered features in the wild, and from the journey of building an AI-driven company itself. Beyond the product, I love connecting with people. I co-founded the AI-Dev community in Israel, dedicated to accelerating innovation in AI coding and product development.