BlogQuality
AI code review when coding agents write more of your code
When coding agents open more pull requests than your engineers can read closely, review becomes the constraint. This guide shows how to let automated reviewers take the first pass and how to keep a named person accountable for every merge.
Two meanings of AI code review
The term AI code review covers two jobs. In one, a model reviews a pull request and posts findings, as CodeRabbit, Qodo, GitHub Copilot code review and Bugbot do. In the other, people review code that a coding agent wrote, which is what most teams mean by reviewing AI-generated code. Teams that adopt coding agents need both.
The pressure comes from volume. An engineer who hands three tasks to agents in the morning can have three pull requests waiting by lunch, and the reviewer may be the first person to read each one closely. Automated review can take over much of the line-by-line checking, but it cannot decide whether a change should exist or answer for it after merge. In the setups we build, the AI reviewer is the first pass and a named person decides.
What automated code review is good at and what it misses
Current AI reviewers read the diff with the surrounding code, and some analyse the whole repository first. They are good at defects visible in the code: unhandled errors and null cases, off-by-one logic, leaked resources, unsafe input handling, secrets in source, tests that assert nothing and breaches of your written rules. They give the fortieth pull request of the day the same attention as the first.
What they miss is mostly knowledge that is not in the repository. These points still need a person who knows the task and the systems around the code.
- Intent: whether the change does what the task asked for
- Architecture: whether the change belongs in this module or duplicates something the codebase already does
- Product correctness: behaviour that is consistent in code but wrong for users, prices, permissions or regulation
- Effects on other services: API contracts, message formats, shared configuration and migrations on live data
- Operational fit: rollout order, feature flags, monitoring and rollback
Coverage also has gaps. GitHub documents that Copilot code review skips dependency management files such as package.json and Gemfile.lock, as well as log and SVG files, so a new dependency needs another check. Check your own tool's exclusions too.
Agent-written changes also repeat patterns that your rules should name. Common ones are tests edited until they pass, broad exception handling, duplicated helpers and edits outside the task.
Setting up an AI reviewer
Most of these tools connect to your git host as an app or run in CI. The work is deciding what the reviewer checks and what happens when it finds something.
- 1
Choose the scope
Start with one or two repositories that receive a steady flow of agent pull requests, and exclude generated code, lockfiles and vendored dependencies. Most tools let you choose when reviews run, such as on opening, on every push or on request, and reviewing every push usually costs the most.
- 2
Write short review rules
Give the reviewer the rules a senior engineer in that repository would apply, starting with what counts as serious here and what to ignore. Leave formatting, linting and type errors to CI. Long rule files dilute the rules that matter, so add a rule only when a real miss justifies it.
- 3
Decide which findings block
Several reviewers are advisory by default, which is a sensible start. If a class of finding should stop a merge, such as a missing authorisation check, enforce it through a required check you control. Requiring the reviewer's own status is not enough, because GitHub treats a neutral result as passing, and Claude Code's review check and Bugbot's default findings both report neutral.
- 4
Control noise and false positives
Cap minor findings per review and suppress new ones on re-review, so a one-line fix does not start another round. For the first month, label a weekly sample of findings as useful, wrong or noise, then adjust the rules and keep the labels as a baseline. Rules that tools learn from your team's replies need an owner who prunes them.
# Review rules: payments-service ## Treat as blocking - Endpoint added or changed without an authorisation check - Query not scoped to the caller's tenant - Migration that cannot be rolled back - Test changed to match new behaviour without a stated reason ## Do not report - Formatting, lint and type errors (CI enforces them) - Generated code under src/gen/ and lockfiles - More than five minor findings per review
Copilot code review reads its instructions from the pull request's head branch, so an agent's change can edit the rules for its own review, while Bugbot reads its configuration file from the base branch. Keep review rule files under CODEOWNERS and treat changes to them like changes to CI.
Review inside agent workflows
The cheapest finding is made before a pull request exists. An agent workflow can add a review step in which a second agent with a fresh context, or a different model, checks the diff against the task without the assumptions the authoring agent carries. Claude Code's local /code-review command, for example, runs as a background subagent with its own context window.
The pstack plugin for Cursor has an interrogate skill that sends the same diff and rubric to one read-only reviewer per configured model and sorts the findings without changing code, treating agreement between models as the strongest signal. Its shipping playbook lands a pull request only on a verdict from an agent that did not write the code, and states that green CI and an approving bot review are not verdicts.
Runmill, an open-source project by our founder, applies the same rule in a pipeline. The exact candidate commit must pass the required checks and a review in a fresh context before a GitHub pull request is opened, and deterministic code, never the agent, decides pushes, pull requests and merges. It is a developer preview, and automatic merge is experimental.
None of this replaces the human review at the end. Skills, playbooks and review prompts instruct an agent but enforce nothing; branch protection, required checks and the agent's permissions do.
Guide to pstack workflows in CursorAdopting the pstack plugin for CursorRunmill source on GitHub
Acceptance standards for agent pull requests
Reviewers move faster when they know what an agent pull request must contain. Write the standard down, send back pull requests that miss it before anyone reads the code, and put the same standard into your agents' instructions.
- One task per pull request, small enough to review in one sitting, with mechanical changes kept apart from behaviour changes
- A link to the issue, spec or prompt the agent worked from
- Tests for new behaviour, and a reason for every changed or deleted test
- A description of what changed, what was left out, what the agent was unsure about and how it was verified
- No unrelated files or new dependencies unless the task asked for them
- The agent and model that produced the change
Agree what the reviewer may assume: that CI built and tested the commit under review, and that AI findings were fixed or dismissed with a reason. The reviewer should not assume that the agent's summary is accurate, that the tests cover the change or that the AI reviewer checked intent and effects on other services.
Keeping a person accountable for every merge
A review comment from a model is input. An approval is a decision that someone answers for, and for agent-written code that someone has to exist. By default, Copilot's reviews do not count toward required approvals, and Claude Code's Code Review neither approves nor blocks. OpenAI's Codex documentation states that review rules do not replace tests, branch protection or required approvals.
Some tools can now approve. Copilot approvals, in public preview, let Copilot submit an approving review that satisfies the required-approval rule once enabled in repository, organisation and enterprise settings, and CodeRabbit's optional request changes workflow approves when its own conditions are met. Decide explicitly whether any automated approval may count toward your merge rules, and if so, allow it only on low-risk paths.
Treat the engineer who delegated a task as its author, and have someone else approve. GitHub builds this into its own agent: whoever asked Copilot cloud agent to create a pull request cannot approve it. With other agents, check whether the person who started the run can approve the result, and if so, close the gap with a team rule or a required check.
Branch protection settings that enforce it
- Require a pull request with at least one approving review before merging
- Require review from code owners, and name owners for agent configuration and review rules in CODEOWNERS
- Dismiss stale approvals when new commits change the diff
- Require approval of the most recent push by someone other than the person who pushed it
- Require build and test status checks to pass before merging
Measuring review burden without scoring people
If agents add pull requests faster than review capacity grows, either the queue gets longer or review gets shallower, and only data tells you which. Measure the review system by team, repository, task type and tool, against a baseline from before agents or a new reviewer arrived.
- Time to first human review after a pull request is marked ready
- Human review time per change, read against diff size
- Comments and review rounds per change
- Share of AI findings acted on: the flagged code changed or the thread was resolved with a fix
- Rework and reverts of merged changes within an agreed window
- Escaped defects traced to merged changes, split by agent-written and human-written work
The share of findings acted on is the most direct measure of noise. Compute it from pull request history, because reactions on review comments depend on people remembering to click them. Label agent pull requests consistently so every metric can be split by who or what wrote the change.
Keep these numbers out of individual performance reviews. Once a metric rates engineers, behaviour adapts to it: changes get split to look fast, and approvals come before the reading. Report by team, task type and tool.
The main AI code review tools
The descriptions below follow each vendor's documentation as of October 2026, in no particular order. Check the current documentation before you decide.
Dedicated review products
CodeRabbit reviews pull requests and is configured through a .coderabbit.yaml file or its web interface. By default it uses agent instruction files such as AGENTS.md and CLAUDE.md as review criteria, and an optional request changes workflow can request changes and approve.
Qodo runs a multi-agent review in pull requests, in which a judge agent merges findings, removes duplicates and drops low-confidence ones. It builds review standards from your codebase, pull request history and requirements, and its Rule Miner turns recurring review patterns into enforced rules.
Review in coding agent platforms
GitHub Copilot code review works on GitHub pull requests, on request or automatically, and can analyse the whole repository for context. It reads repository custom instructions, path-specific instruction files and AGENTS.md, and its approvals feature is in public preview.
Bugbot is Cursor's pull request reviewer for GitHub, GitLab, Bitbucket and Azure DevOps. Its rules live in .cursor/BUGBOT.md files, and Cursor's general project rules do not apply to it. An optional Autofix starts a Cursor cloud agent to fix what it found.
Claude Code offers a managed Code Review for GitHub, in research preview for Team and Enterprise subscriptions. Several agents review the diff in parallel, a verification step filters false positives, and a REVIEW.md file tunes what is flagged. Teams can also run Claude in their own GitHub Actions or GitLab CI/CD pipelines.
OpenAI's Codex code review posts a standard GitHub review when someone comments @codex review, or automatically once enabled. On GitHub it flags only P0 and P1 issues, and it applies review rules from AGENTS.md files to the files they cover.
When choosing, compare where code is sent and under which terms, which instruction files each tool reads and how its approvals and checks fit your merge rules. Then test the candidates on your own merged changes.
How Cloudsail helps
We evaluate review setups on your own merged changes. We run past pull requests, including ones that later needed a fix or a revert, through the candidate reviewers and configurations and compare the findings with what actually went wrong. You see which defects each setup would have caught, and how much noise and cost it adds per reviewed change.
We write the review rules with your engineers, along with the repository instructions, CODEOWNERS entries and branch protection they depend on, each with an owner. We also set up the measurement of review burden at team and tool level.
Reviewer training runs as team sessions on your own repositories, remotely or on-site in Germany and Poland, in English, German or Polish. It focuses on reading agent pull requests: checking intent against the task, spotting patterns agents repeat and deciding when to send a change back. We do not run public courses or issue certificates.
Evaluations on your repositoriesHarness engineeringGitHub Copilot implementation
Questions
Can an AI reviewer replace human code review?
Not for changes that matter. It can take over much of the line-by-line checking and run on every push, but it cannot judge intent, architecture or product correctness, and it cannot answer for a merge. We use it as the first pass and keep a required human approval.
Should AI review findings block merges?
Only narrow, well-defined classes of finding, such as a missing authorisation check, and only through a required check you control. Some tools report findings as neutral by default and GitHub counts neutral as passing, so requiring their check alone does not block on findings.
Which AI code review tool should we use?
We do not rank tools, because the answer depends on your git host, the instruction files you already maintain, your data constraints and your merge rules. We compare candidates on your own merged pull requests and show what each would have caught and how much noise it adds.
Can the engineer who started an agent approve its pull request?
We recommend treating that engineer as the author, so someone else approves. GitHub enforces this for Copilot cloud agent. With other agents, check who the pull request is attributed to and add a rule or a required check where needed.
Does an AI reviewer send our code to a vendor?
Hosted reviewers read your code on the vendor's infrastructure and under the vendor's terms. Agree data handling and retention with your security team before you enable one, and check availability for your contract. Claude Code's managed Code Review, for example, is not available to organisations with Zero Data Retention enabled.
Related services
Evaluations on your repositories
Tools, models and settings compared on work your team has already merged.
Harness engineering
The instructions, commands, permissions and checks that decide how an agent behaves in your repositories.
GitHub Copilot implementation
Copilot and its cloud agent, working within the policies and review process you already have.
More from the blog
Coding agent security
Practical security and governance for coding agents: prompt injection, secrets, permissions, sandboxing, unattended runs and ownership.
pstack workflows
What the pstack plugin for Cursor contains, how its playbooks and model roles work, and how a team adopts it under its own controls.
Autonomous software factory
What it takes for coding agents to carry scoped work from issue to merge, which work fits, and where people stay in control.
Sources
- Claude Code documentation: Code Review
- GitHub Docs: About GitHub Copilot code review
- GitHub Docs: Risks and mitigations for GitHub Copilot cloud agent
- GitHub Docs: About protected branches
- Cursor documentation: Bugbot
- OpenAI: Review GitHub pull requests with Codex
- CodeRabbit documentation: Request changes workflow
- CodeRabbit documentation: Code guidelines
- Qodo documentation: The Qodo Code Review experience
- pstack plugin for Cursor: README and skills in cursor/plugins
Talk to an engineer
A 30-minute call with one of our engineers about your coding agent setup, what it costs and where it can improve. No access to your systems and no data shared.
Not using coding agents yet? Use the same form and tell us what you are planning.
Our team has built software and production AI for trivago, SAP, Tonies, EWE and tecRacer.
Check your email program
We tried to open a draft in your email program. Nothing is sent until you send it, and this page cannot tell whether the draft opened or the email was delivered. The form stays editable, and you can also write to miki@cloudsail.com directly.