AI agent code review starts with a decision made before the agent runs: whether its output lands as a diff a person can read, or as a write nobody sees. A silent write modifies the working tree, or commits and pushes, with no step between the agent finishing and the change being real. A reviewable diff is a proposed change with a base, a boundary and an approver.
Everything else in this post follows from preferring the second. It is not a statement about how good agents are. It is a statement about where the last gate in a software process should sit, and that it should not move because the author changed.
Silent writes versus reviewable diffs
A diff has three properties a silent write lacks. It has a base: the exact state it was computed against, so a reviewer knows what the before was. It has a boundary: the set of files and lines touched, which is the first thing a reviewer scans. And it can be rejected as a unit, which is what makes saying no cheap.
Teams already have the machinery: branches, pull requests, required checks, branch protection, and the git diff format all of it is built on. An agent that lands its work as a pull request plugs into habits that exist. An agent that writes to a shared branch asks the team to invent new ones under pressure.
There is a cost: the agent's work is not done when the agent stops. Review time is real, and it is a cost the agent's own metrics never show, which is why it belongs in the numbers when you evaluate a coding agent on your own repository.
Blast radius and attribution
Blast radius is what a wrong change can reach before anyone notices. A diff bounds it in space, to the files touched. Review bounds it in time, to before merge. Permissions in the agent harness bound it in kind: no deploy credentials, no write access outside the repository, no network beyond what the task needs.
Attribution is the other half. For every agent-authored change, someone should be able to answer: who asked for it, which model and version produced it, what tools and commands ran, and who approved it. Put the first three in the commit or pull request body, and keep the full trace linked from it.
This matters for audit, and it matters for cost. Per-team LLM cost attribution has to be attached at call time; an agent that does not record which developer and task a run belonged to produces a bill that cannot be explained either. Runix Code is designed with this as a team control: usage attributed by developer, alongside seats and per-key limits.
Tests as the gate
The agent's claim that tests pass is not evidence. Continuous integration running the tests against the diff is. Two kinds of test matter for an agent's change, and they are the same two that code tasks in Runix Data are verified with: the tests that should fail before the change and pass after it, and the rest of the suite, which should pass both times.
Agents will, on occasion, make a failing test pass by editing the test. Policy follows from that: any diff that modifies tests gets its test changes read first and separately, and a change that deletes or weakens an assertion is rejected by default unless the task asked for it.
If the task did not come with a test, the first thing to review is whether the agent wrote one and whether that test would have failed on the old code. A change without a failing test behind it is a change whose correctness is being asserted, not checked.
How AI agent code review differs from reviewing a human
The standard code review practices still apply: design, functionality, complexity, tests, naming, comments. What changes is the prior. A colleague's diff comes with shared context, a conversation and a reputation. An agent's diff has none of those, and it has specific failure modes a person rarely shows.
- Plausible, not correct. Agent code reads fluently. Fluency is not a signal; run it in your head against an edge case the tests do not cover.
- Over-scoped. Formatting changes, renamed variables and refactors nobody asked for. Each one hides the real change. Reject on scope before reviewing on content.
- New dependencies. An added package is a supply-chain decision. Treat it as a separate review.
- Confident comments and commit messages. They describe intent, not what the code does. Read the code.
- Silent behaviour changes. A changed default, a swallowed exception, a widened type. Diff the public surface, not just the lines.
- Edited tests. Covered above; read them first.
Two habits follow. Read in the order task, tests, diff: know what was asked, check what was proven, then read what was done. And when a change is wrong, reject it and re-run the agent with a sharper task rather than fixing the diff by hand, so the trace stays honest and the next run gets the correction. Small diffs are not a style preference here; they are what makes that order workable.
What to log
The review record for an agent-authored change should let a later reader reconstruct it without the agent. Per change:
- The task as given, verbatim, and who gave it.
- Model id and version, and the harness version.
- Every tool call and shell command, with its exit status.
- Test runs: which tests, before and after, with results.
- Tokens, cost and wall-clock time for the run.
- The diff hash, the reviewer, and the decision with its reason.
Prompts and file contents in the trace fall under the same retention and redaction rules as any model traffic; what to log for LLM traffic covers which fields to keep and which to strip. The trace is also what makes a run replayable, which is the only way to debug an agent that produced the wrong change for reasons nobody can see in the diff.
Runix Code is in development and is designed around this gate: every change lands as a reviewable diff, never a silent write, with team controls for seats, per-key limits and usage attribution by developer. The rollout guide describes the review gate as the spine of the product rather than a setting, and the Runix Code page states what is being built and its current status.
Questions this raises
Why should an AI agent's changes land as a pull request instead of a direct commit?
A pull request has a base, a boundary and an approver, so a wrong change can be rejected as a unit before it is real. It also reuses the review machinery the team already has rather than requiring new habits.
Should I trust an agent's report that the tests pass?
No. Run the tests in CI against the diff. Also read any test changes first, because an agent can make a failing test pass by weakening it.
What should be recorded for an agent-authored change?
The task and who gave it, the model and harness versions, every tool call and command, test results before and after, tokens and cost, and the reviewer's decision. The full trace should be linked from the pull request.