Transcript
In this episode
Honeycomb went from 30 to 70 merged pull requests a day in three months. The catch: automated code review, not code generation, became the real bottleneck of software automation. Liz Fong-Jones, Technical Fellow at Honeycomb, explains why incidents still rose 1.5x, how their internal bot Autobot now reviews every PR, and why AI amplifies whatever org you already have.
What we cover:
- How automated code review lets humans focus on design, not trivial bugs
- Using a decision model like Jev to decide which PRs are safe to auto-merge
- What makes a codebase ready for AI coding agents
- When to trust AI agents with production incidents, and when they're just throwing darts
- Why "Claude did it" isn't an excuse, and what ownership means with AI
- How open source maintainers can handle a flood of AI slop pull requests
Chapters:
00:00:00 - Introduction
00:06:41 - Why AI amplifies dysfunctional engineering orgs
00:10:45 - What makes a codebase AI ready
00:14:01 - Honeycomb's Autobot and automated code review
00:19:06 - Using Jev to decide which PRs are safe to merge
00:26:51 - Trusting AI agents during production incidents
00:30:40 - Least privilege and guardrails for coding agents
00:33:25 - If your name's on it, you own it
00:38:55 - AI slop pull requests and open source
00:45:56 - Will observability engineering survive as a role?
Build your software factory, one workflow at a time, with Tessl:
https://tessl.co/nlr
π Subscribe for weekly episodes on AI-native development
Is your team's review capacity keeping up with your AI coding agents? Tell us in the comments.
Automated Code Review at 70 PRs a Day: Lessons From Honeycomb
When a team doubles its output with AI coding agents, the first thing to break is rarely the code itself. It is the humans' ability to check it. At Honeycomb, merged pull requests went from roughly 30 a day to 70 in three months, and the lesson that emerged is that software automation only pays off when automated code review scales alongside generation.
Liz Fong-Jones, Technical Fellow at Honeycomb, has spent her career in reliability, from SRE and managing the Bigtable team at Google to DevRel and field CTO work. On The AI Native Dev, she joined host Simon Maple to unpack what Honeycomb learned, why AI tends to magnify whatever organization it lands in, and where humans still need to hold the wheel.
AI Amplifies the Org You Already Have
Simon opened with one of Liz's own lines: AI makes a dysfunctional org more dysfunctional and a high-ownership org faster. Rather than cataloguing dysfunction, Liz anchored on the positive trait that seems to predict everything else. "I think ownership is really that primary force that dictates whether an org is high or low functioning," she explained.
To make sense of this, it can be helpful to plot teams across two dimensions: ownership and production maturity. High ownership without mature tooling produces conscientious engineers who are slowed down by five sets of permissions. Mature tooling without ownership looks more like the scenario Liz pointed to at Shopify, where an AI-first mandate was followed a year later by AI slop landing on teams. Her read was that it was a foreseeable consequence of telling people to move faster without owning the results. Only teams with both appear to benefit cleanly. Without direction, as she put it, rocket fuel just sends a rocket in circles.
The practical implication is simple: AI will not fix an existing problem, it will make that problem arrive sooner.
What Makes a Codebase Ready for AI Coding Agents
Liz's first answer was about patterns. "If you have five different ways of doing something, it will become very confused and it will create six ways or seven ways of doing it," she noted. Agents copy what they find, so alignment on common patterns, encoded for both agents and new hires, becomes a prerequisite rather than a nice-to-have.
The rest of the list will sound familiar to anyone who has worked in reliability: fast CI/CD that validates a change in five minutes rather than three hours, meaningful tests, and good observability. Even Honeycomb, with relatively high-quality code comments, found AI regressing it toward the mean.
The twist is that the vegetables are now easier to eat. Because machines never get bored, adding telemetry, rich wide events and attributes to a code base is exactly the kind of work agents handle well, as long as someone tells them to do it. In practice, both the necessity and the ease of this work have gone up.
Automated Code Review Is the Real Bottleneck
Honeycomb's internal bot, Autobot, began life as Claude Code in a box, turning a well-decomposed Linear ticket into a pull request. Liz pointed out that it only produces around five PRs a day. It now acts as the reviewer on every pull request, and it has become good enough that developers no longer hunt for trivial bugs themselves.
The pivot came from economics. Anthropic's Claude Code review functionality impressed the team, but at 20 to 30 dollars and 10 to 20 minutes per run, it did not fit their volume. They aimed instead for something that ran in two or three minutes and cost dollars per pull request, because, as Liz observed, human developers were already more than capable of creating mass volumes of PRs, with or without Autobot.
The next layer is deciding what humans should look at. Honeycomb had been experimenting with having around 20% of pull requests reviewed and landed automatically, relying on engineers to self-identify low-risk changes. Liz now sees a decision model like Jev acting as the gut check: given the review, can this land safely? The confidence score, more than cost or speed, seems to be what matters.
This is not about reviewing less. "It's to focus the time that we're spending doing code reviews on the design patterns and not on trivial bugs," she argued. The numbers suggest why it matters. Initially the change failure rate looked like it was falling, but over the longer term Honeycomb was shipping twice as many pull requests and seeing 1.5 times as many incidents. There is a limit to how many incidents humans can work.
Trusting AI Agents in Production Depends on the System, Not the Model
Can an agent handle a 2 a.m. incident while you sleep? It depends on how well a team has implemented the practices of the last decade. With working feature flags, automatic rollbacks and observability, an agent with MCP access can flip the right flag and leave a record for the human reviewing it over coffee the next morning. Without them, "your agents are just guessing and throwing darts at the dartboard."
Her design principle is that the tools themselves must be safe. Rather than trusting an agent's judgment about whether to kill pods in production, give it tools to abort or revert rollouts. The same applies to your most junior developer.
She also made a counterintuitive case against one common safeguard. Having humans click yes on every action does not scale, and classifiers appear to do a better job when attention is scarce. The mistake, in her framing, is over-trusting our own ability to make decisions under time pressure and stress.
If Your Name's On It, You Own It
Honeycomb's AI values state that if your name is on it, you own it. Liz was candid about her own lapse. Her AI assistant failed to find a partition ID field, confidently suggested creating one, and she approved a pull request that a principal engineer then spent time reviewing. The field had been there all along. Blaming Claude is not an excuse; the more useful question is which structural biases led to over-trusting the output.
Even in fully automated loops, a human team still owns the workflow files, the system prompt and the quota, and remains responsible for spot checks and evals.
Open Source and the Future of Platform Engineering
On open source, Liz suggested maintainers should keep welcoming well-formed bug reports while treating pull requests as discretionary, since their own agent can often produce a fix in their preferred style. Quizzes about submitted work, and Chrome security's "no crasher, no triage" approach, show how AI can help filter AI slop. The unresolved question is who pays for the tokens that keep maintainers from burning out.
Looking ahead, she sees observability engineering continuing its merger into platform engineering alongside security, testing and reliability. You don't want every developer to be an expert in service level objectives, but you do want every developer to have them.
The full conversation is worth a listen for any team whose AI coding agents are outpacing its review capacity. How is your team deciding what still needs a human reviewer?
CHAPTERS
