ARTICLE
Anatomy of an Agentic Code Review
Explore the agentic code review process and how it adapts to fast-paced PR environments. Learn to optimize your reviews with Tessl Code Review.

Simon Maple

The path from raising a ticket to merging a PR has looked the same for most teams for many years now. Once a ticket lands in Linear, or Jira, a developer would sit down at their mechanical keyboard of choice, write the code and open a pull request. A colleague would review it (spending an inversely proportional amount of time to the number of lines of code) and after a round or two, it'll be merged. Review was never fast, but it kept pace because humans wrote code at human speed.
Coding agents abruptly changed that reality. They can turn a ticket into a pull request in minutes, so teams end up getting greater numbers of PRs, much faster. Review didn't speed up with them and, as we well know, it became the bottleneck.
This post looks at the new world of code review built for a world where PRs are being raised at increasing rates. Here's the whole agentic workflow we'll walk through. Through the rest of this post, we'll build this diagram up one part at a time, starting from the traditional review it replaces.
TL;DR
- A typical code review tends to be automated tests and human review, but when we involve agents we can add multiple new angles such as verifying coding standards, review persona perspectives, feedback and remembering past mistakes and more.
- Beyond tests, most human reviews often only add a few loosely stated findings that vary from person to person. At agentic PR volume reviews slow down, vary by reviewer and forget what they learned.
- If you store knowledge and practices in context, code review is a great stepping stone towards a context-driven software factory.
Get started: interested in doing this yourself? You can, and it’s free to get going, simply install Tessl Code Review on your repository and see the first loop running on your next pull request. Set up Tessl Code Review.
The review you're replacing
Let’s start by refreshing our memory on the traditional flow of a code review. A developer raises a PR and a human takes a look, offers feedback and guidance, and a manual loop begins till the PR changes are approved and merged.
A typical human review on this pull request might be as follows:
“Took a look, mostly looks good. Have you thought about what happens if the user ID is missing? There's also already a retry helper in utils, might be worth reusing. Happy to approve once that's sorted.”
There's nothing wrong with this comment. It's friendly and the points are fair. But it doesn't say which issue blocks the merge and which is a suggestion. It doesn't point at any lines of code. It mixes a correctness question with a code-reuse nudge. And it misses the most serious problem in the change. Multiply that by the volume an agent produces and no human reviewer can keep up, let alone stay consistent. Of course, the quality of the review will vary depending on which human reviewer in the team was able to pick it up.
The anatomy of a agentic code review
- A standard (defined before): coding standards, best practices, what good should look like in the codebase, and design choices.
- Perspectives (defined before): the personas or types of expert opinions should vary, and change reviews are examined from several angles, each asking one narrow question.
- Findings (discovered during): each issue located, weighted and explained.
- A verdict (decided during): a more detailed decision that follows from the findings.
- A response to every finding (built after): an actionable next step: fix, refute or decline, with the next round respecting the answer.
- A memory (updated after): what the review learned, fed back into the standard so the next change review respects past decisions and learnings.
The human comment we shared earlier has a couple of loosely stated findings and not much else. Let's take the same pull request but this time, we’ll run through each of these parts in turn.
1. A standard
Every review checks a change against something. In the comment above, that something is whatever the reviewer happened to know and care about that day. A different teammate would have checked against a different standard, and the author had no way to know either one in advance.
A good review starts from a standard that's written down and visible before the pull request is opened. The author can check their work against it, the reviewer applies it consistently and the team can argue about the standard itself rather than about individual comments. In Tessl Code Review, that standard lives in your repository as context: the skills, docs and rules your team owns and changes like any other code.
2. Dedicated perspectives
A single reviewer reading a diff top to bottom is juggling correctness, security, performance and readability at once. Whatever catches their eye gets the attention. Here, that was a missing user ID and a duplicated helper, while the most serious problem went unnoticed.
A good review splits the job up, and gives each perspective one narrow question to answer:
- Correctness and data integrity: does the change do what it's meant to, and does the data it touches survive it intact?
- Security and privacy: what could an adversary do with this change, and what does it expose about people?
- Scale and resilience: what does it cost as load grows, and what happens when something it depends on fails?
- Maintainability and code quality: can the next person or agent understand it and change it safely?
In Tessl Code Review, each perspective is a lens: a focused review perspective, packaged as a skill, that applies your standard to one area. These four are the built-in lenses, and they run in parallel. Asked only what an adversary could do, the security perspective spots what the generalist missed: the endpoint returns email addresses to any signed-in user. You can of course add to these if you wish and define your own lenses which make sense for your organisation and projects.
3. Findings
Here's the same pull request, reviewed perspective by perspective:
Changes required
- Required · Security and Privacy: returns email addresses to any signed-in user.
src/api/users.ts:48 - Advisory · Maintainability and Code Quality: rebuilds the existing
retry()helper. - Advisory · Scale and Resilience: no timeout on the payments API call.
Each finding has the same structure:
- Where it came from. The perspective that raised it, so you know what kind of problem it is before you read it.
- How much it matters. Required or advisory, decided by the reviewer rather than left for the author to guess.
- Where it is. A file and line, so nobody has to work out what the reviewer meant.
- Why it's a problem. The consequence, stated plainly. “Returns email addresses to any signed-in user” tells you the risk, not just that something looks off.
Compare that with “mostly looks good”. Every finding here can be acted on without a follow-up conversation.
4. A verdict
A list of findings isn’t a decision. A good review ends with one: approve, or request changes. Some findings may be required and some may be advisory, and it’s often easier to relate to this when thinking about security issues being raised. That verdict should follow based on the triage result. The two advisory findings are worth fixing, but on their own they wouldn’t stop the merge.
The bar for what blocks should be explicit too. In Tessl Code Review, the team sets the minimum severity that can request changes, so “blocking” means the same thing on every pull request. The human reviewer is still part of the process, but they start from a structured verdict instead of a blank diff, and spend their time on judgement calls rather than spotting the obvious.
5. A response to every finding
A review is a conversation, not a broadcast. A good one expects a response to every finding: fix it, refute it, or decline it with a reason. Declines are as useful as fixes. “Intentional, the caller already applies a timeout” is knowledge the next review should have.
The next round should respect those responses. When Tessl Code Review re-reviews a pull request, findings that were fixed are marked resolved, findings still present aren’t posted again and findings that come back after being fixed are reopened. The author never has to wade through the same comment twice.
Beyond the code review: Tessl verifiers and security findings
Before we get to the final part, memory, it's worth looking at what a good review shouldn't have to catch at all. Two kinds of problem are better handled either side of it.
Tessl Verifiers, before the pull request
A verifier can a deterministically triggered based on the type of change that is made. For example if you update some front end code, you might have a front end verifier that then uses an LLM to check and enforce a specific rule on the committed code. Verifiers live in your repository alongside the rest of your context, and can be run before the pull request is opened as well as during normal review in CI. If a verifier fails as code is being written, the agent fixes the problem before the review stage in CI. Verifiers pass or fail in a binary fashion shown below:
1✓ imports-use-path-aliases
2✓ no-secrets-in-source
3✓ migrations-are-reversible
4✗ api-routes-require-auth src/api/users.ts:48They're needed because judgement is expensive. Every finding a lens raises costs a review, a round trip back to the author and another review. A verifier runs in seconds, costs almost nothing and gives the same answer every time. Anything a verifier can catch, review shouldn't have to, which leaves the lenses free to spend their attention on problems that genuinely need reasoning about.
Verifiers aren't written once and left alone, and review feedback is how they evolve. When a lens keeps raising the same finding and the rule behind it turns out to be binary, that's a new verifier waiting to be written. The failure above is exactly that: the auth gap the security lens flagged at src/api/users.ts:48, now caught before the pull request exists. Feedback works the other way too. If a verifier fires on code your team decides is fine, the decline is a signal to tighten or retire the check. Either way, the set of verifiers grows and sharpens with every review, which is the memory we'll come to next.
Security findings, alongside the review
Your security scanner, whether that’s Snyk, Semgrep or Dependabot, checks the change for known vulnerabilities. Its findings are presented in the same way as review findings are so they can be sent straight back to the coding agent to fix, then scanned again, rather than left in a dashboard waiting for someone to triage them.
- High · Dependency: lodash 4.17.15 has a known prototype pollution issue. Upgrade to 4.17.21.
package.json - Medium · Code: search query built by string concatenation. Use a parameterised query.
src/db/search.ts:22
A remediation policy in your repository decides what gets fixed automatically. Together, verifiers and security findings take a whole class of problems off review’s plate, which leaves it free to spend its attention on the problems that genuinely need judgement.
6. A memory
Possibly the most important part of the longer term process. We don’t want agents to make the same mistakes over and over, which is incredibly frustrating and makes us lose trust in it’s ability. Most reviews end at merge, and everything they learned leaves with them. The next pull request makes the same mistake, and the same review comment appears again.
A good review leaves something behind. Look at our three findings from our PR and what each one should become:
- The auth gap. If review keeps finding unauthenticated routes, that’s not a judgement call any more. It becomes a verifier, so the problem is caught before the pull request even exists.
- The rebuilt retry helper. If agents keep rebuilding it, they don’t know it exists. That belongs in the context the agent reads before it writes code.
- The missing timeout. If it’s a standard, the scale and resilience perspective should check for it explicitly every time.
Each finding stops being a comment someone has to repeat and becomes a fix at the source.
This has become very powerful to us at Tessl. We have build a maintenance agent that periodically reads historic pull requests, findings and logs, then proposes changes like as pull requests that our team reviews.
The anatomy, in one line
While speed of review is a major way of how agentic code review can improve our workflow, there are many more important aspects that give us a more reliable, trusted process. Using context that reflect our best practices and codings standards that review agents use, as well as verifiers, and lenses give us the outcomes we most want from a review. Further to this, the learning stage allows us to improve the context with every PR, using human direction, and comments to avoid making the same mistakes time and time again, for each PR.
That’s why code review is the right place to start building context. When the standard lives in your repository and every finding feeds back into it, review becomes a stepping stone to a context-driven software factory, where every change starts from what the last one taught you.
Get started: install Tessl Code Review, then generate your first lens from your existing code and past pull requests. Set up Tessl Code Review
COPY & SHARE

Simon Maple
Simon Maple is the Head of Developer Relations at Tessl, and AI Native Dev co-host. Previously, Simon was the Field CTO, and VP Developer Relations at Snyk, ZeroTurnaround, and IBM. He became a Java Champion in 2014, JavaOne Rockstar speaker in 2014 and 2017, Duke’s Choice award winner, Virtual JUG founder and organiser, and London Java Community co-leader.
READING
·
0%
IN THIS POST
COPY & SHARE

Simon Maple
Simon Maple is the Head of Developer Relations at Tessl, and AI Native Dev co-host. Previously, Simon was the Field CTO, and VP Developer Relations at Snyk, ZeroTurnaround, and IBM. He became a Java Champion in 2014, JavaOne Rockstar speaker in 2014 and 2017, Duke’s Choice award winner, Virtual JUG founder and organiser, and London Java Community co-leader.
YOUR NEXT READ
A context driven code review that learns and improves on every run
Tessl's context-driven code review integrates with software development, improving code generation and maintenance by learning from past reviews and enhancing standards.

Simon Maple



