Review Agent-Written Code Against Your Team's StandardsLearn More
Book a Demo
CareersDocs
Log inBook a Demo

PODCAST EPISODE 126

Jacob Lauritzen: Token Maxing Is the New Lines of Code

AI writes almost all of Legora's code. So why does its CTO think token maxing is a mistake, and why should litigation strategy stay a human call?

6 Oct 202644 min 21 sec

Transcript

In this episode

AI coding agents now write almost all of Legora's code, with engineers running up to ten in parallel. Yet the CTO behind that shift thinks measuring adoption by tokens burned is a mistake. Jacob Lauritzen, CTO of Legora, joins Simon Maple to unpack what actually scales when agents do the typing, and where human judgment still has to stay in charge.

What we cover:

  • How Legora's engineers orchestrate many AI coding agents at once, from shared local dev resources to cloud agent setups
  • An agentic loop that reproduces, fixes and tests bugs straight from Slack
  • Why token maxing is a bad way to drive AI adoption, and what to measure instead
  • Agent evaluation: regression testing the harness as models and prompts change every week
  • Decision models, the verifier's rule, and where autonomy ends and human judgment begins
  • Why AI agents need more than a chat interface

Chapters:

00:00:00 - Introduction

00:01:47 - Legora's growth from $100M to $200M ARR

00:03:44 - Parallel coding agents and bug-fix loops

00:07:15 - Senior engineers vs AI-first engineers

00:10:10 - Why token maxing is the new lines of code

00:14:13 - What AI can and can't verify in legal work

00:17:30 - Decision models vs autoregressive LLMs

00:24:22 - Weekly model benchmarks and harness evals

00:26:58 - Why AI agents need more than a chat interface

00:35:03 - The verifier's rule and litigation strategy

Build your software factory, one workflow at a time, with Tessl:

https://tessl.co/lpg

๐Ÿ”” Subscribe for weekly episodes on AI-native development

Is token count telling you anything real about how your team uses AI coding agents? Tell us in the comments.

Few engineering teams have leaned into AI coding agents as hard as Legora's. The legal AI company, which went from $100 million to $200 million in ARR in six months, now has AI writing almost all of its code, with engineers orchestrating many agents at once. That makes it a useful place to ask a harder question: once agents are doing the writing, what should leaders actually measure, and where does human judgment still belong?

Jacob Lauritzen, CTO of Legora, joined Simon Maple on The AI Native Dev to work through both. His answers suggest that the teams getting the most from agents may be the ones paying the least attention to how much AI they appear to be using.

Orchestrating many coding agents at once

Running a handful of agents in parallel is less a matter of willpower than of infrastructure. Lauritzen described a stack of internal tooling built to make it possible: a custom background coding agent, cloud development environments where each agent gets its own logs and observability and can send back screenshots and video of its work, and an engineer-built Kanban board for agents.

Local development needed rethinking too. Legora's app depends on Postgres, Redis and an observability stack, and spinning up a full copy per worktree would quickly exhaust a laptop. The team built a tool that runs one instance of each dependency and shares it across workspaces, giving every agent its own database and prefix instead.

The more striking example sits outside any developer's machine. A bug reported in Legora's Slack channel kicks off a cloud agent that tries to reproduce it. If it succeeds, it posts video evidence back to Slack, fixes the issue, writes a regression test and opens a PR for the on-call "goalie" engineer to review. In practice, this is an agentic loop where the human arrives at the end rather than the beginning.

Why token maxing is the wrong adoption metric

As usage grows, leaders want to know whether their organisation has really adopted AI. For a while, the popular answer was token maxing: leaderboards of token consumption, managers telling engineers they were not using enough, and in some companies, token usage written into career ladders.

Lauritzen never ran that playbook, and he pointed out what happened elsewhere: engineers "that would just burn tokens to burn tokens, right? They just have them running in loops, doing nothing." When Simon called it the new lines of code, Lauritzen agreed without hesitation.

His alternative is simpler. "The thing that's most important to me is: are the teams delivering great stuff really quickly?" he explained. If a team hit that bar without AI, he would be fine with it; he simply does not think it is possible. Adoption then follows from high expectations, visible wins and hackathons that open people's eyes, rather than from a quota.

That does not mean nothing is measured. Legora's developer experience and platform teams look at efficiency signals such as how many turns it takes an agent to reach the desired result. The framing is organisational leverage: investing in AI enablement and guardrails as a single knob that can make every engineer 5 or 10 percent more effective, something leaders have wanted for years but could rarely deliver so quickly.

Agent evaluation: testing the harness, not just the model

Legora benchmarks many models every week and switches between them per task, decomposing work across sub-agents so each piece runs on whatever is Pareto efficient. Lauritzen was clear that not every task needs frontier intelligence: "Some of it just requires a completely ordinary intelligence. But if it's very fast and cheap, that's great."

That pace of change raises the stakes on evaluation. Asked whether robust, "torture test" style suites still matter, Lauritzen noted that the harness changes as often as the model does. The team might alter task decomposition, compaction, file search, prompts or tool search. Every hard case or failure gets added to an extensive eval set as a regression test, and the suite runs on every change so the agent keeps performing at its best.

The verifier's rule and the limits of autonomy

It can be helpful to plot agent work along a single dimension: how cheaply can the answer be checked? Lauritzen's framing starts from the verifier's rule.

What is the verifier's rule? The verifier's rule holds that AI will eventually master any task with an objective answer that is quick and cheap to check, even when producing that answer is hard. Maths, and code backed by tests, are the classic cases. Put a model in a harness that tells it when it is wrong, and it tends to converge on the right answer.

Law is full of tasks that fail that test. Ask five partners at five firms for a litigation strategy and you may get five different answers. Simon pressed the obvious counterpoint: if the humans disagree anyway, why not let the AI pick one? Lauritzen conceded that a client could, but argued the gap is context rather than capability. A client's lawyers know their risk appetite, their history and where the business is heading in five years. "If there's no objective truth, that should be a human who makes that judgment," he argued, at least until that context can flow into the system.

Responsibility follows the same logic. "I'm still responsible for my PR that I put up, even though I did it with Claude, or even though Cursor wrote all the code," Lauritzen observed, and lawyers using Legora own their output in the same way.

Why AI agents need more than a chat interface

If humans stay in the loop, the interface for that loop matters. Lauritzen argued that a long chat thread reflects a human constraint rather than a natural fit for agents. For a due diligence review across millions of documents, Legora built a spreadsheet-style view with files as rows and data points as columns, backed by multi-hop citations that trace every claim to its sources.

He sees agents moving toward being always on and proactive, coming to the human with a specific question, such as a red flag in a contract and two ways to handle it, rather than waiting for instructions. As trust grows, that threshold seems likely to keep shifting, much as engineers moved from approving every tool call to running coding agents in auto mode.

For engineering leaders, the conversation offers a practical lens on AI coding agents and beyond: measure output, invest in the harness, and design deliberately for the moments where human judgment adds something no test can check. It is a fascinating conversation, and worth a listen in full.

CHAPTERS