An Agent Should Never Merge Its Own Pull Request

Share
An Agent Should Never Merge Its Own Pull Request. Abstract opinion illustration in orange and dark grey on debugly.dev

Here is the thesis, and it is short. An agent that can merge its own pull request has removed the one control in software engineering that does not share the author's blind spots, and no amount of agent self checking puts it back, because the checks are generated by the same model, from the same context, with the same omissions. The merge gate held by a separate approver, human or otherwise independent, is not a tax on autonomy. It is the property that makes agent code reviewable at all.

I am not arguing that agents should not write code, nor that every change needs a slow human ceremony. I am arguing that the author and the approver must be different minds, and that an agent merging its own work is, structurally, a pull request with no approver.

Why self checking is not a second opinion

The temptation is to let the agent review its own diff before merging, and modern agents do this fluently, producing a self review that reads well. But the self review is sampled from the same distribution that produced the code. The model that missed the missing error path in the write is the model that will miss it in the review, because the omission is not a slip, it is a property of what the model considers. A second pass by the same mind is a re read, not a review.

This is the echo chamber in miniature, and at team scale it becomes when an agent reviews an agent, the praise is guaranteed, but even a single agent reviewing itself has the same defect: agreement between two outputs of one model is not evidence, because the outputs are correlated by construction.

The human parallel is exact and familiar. We do not let an engineer merge their own change, not because engineers are untrustworthy, but because the author is the one person guaranteed to share the change's blind spots. The rule is not about honesty. It is about correlation.

The merge is the trust boundary

Everything before the merge is proposal. The merge is the moment the change acquires the authority to affect production, and authority should require an independent judgement, because the independent judgement is what prices the risk the author cannot see. In the agent case the independent judgement also carries the ownership that the workflow otherwise removes, per AI wrote it so nobody owns it: the approver who merges is the owner, and the merge right is precisely what makes them so.

Remove the separate approver and the change enters production with no person who has judged it and no mind outside the author's distribution, which is a strictly weaker position than the weakest human workflow, where at least a tired reviewer is a different brain.

The counterargument, fairly

The honest defence of self merge is speed, and it is real. An agent that waits for a human loses the latency that justified the agent, and for a fleet of agents the human becomes the bottleneck, so teams will be tempted to let the machines merge and sample human review afterwards. Sampling afterwards is not nothing, and for low risk changes it may be enough.

I accept the speed argument for the draft, the experiment and the throwaway branch. I reject it for the merge into the protected trunk, for three reasons. First, the changes that cause incidents are not the ones predicted to be risky, so sampling after the fact misses the dangerous tail. Second, post hoc review discovers the defect after it has had production access, which is the worst time to discover it. Third, the merge gate is cheap and the incident is not, so the asymmetry favours the gate.

There is also the practical point that a separate approver need not be slow. A second, independent agent, from a different model family, reviewing with a different context, is a genuine second opinion in the statistical sense, and can hold the gate at machine speed. The requirement is independence, not humanity.

What the gate should check

If the approver may be a machine, the gate must be designed so that independence is real. The reviewer agent should be a different model, given the diff and the repository but not the author's chain of thought, and asked the disconfirming questions, what does this break, what must it never do, which test would fail if the assumption were wrong, because the adversarial framing is what extracts signal from a correlated mind.

The gate should also enforce the narration, one paragraph on what the change does at the boundaries and how it was verified, per reviewing code the author cannot explain, and reject diffs larger than the narration, because the size to comprehension ratio is the proxy for ownership.

And the gate should log, for every merge, who approved and what they were independent of, so the audit trail can answer the postmortem's first question, which is always "who judged this", and an agent self merge is the answer "no one".

The uncomfortable conclusion

The teams that will get hurt are not the ones using agents, but the ones that removed the separation of author and approver to save a minute, and discovered that the minute was priced in incidents. The separation is old, and it is old because it works, and it works because it is the only control that is structurally uncorrelated with the change.

The agent can draft, test, argue and revise at machine speed, and should. The merge stays a two mind act, because the second mind is the only one that does not share the first mind's omissions, and omitting is exactly what the dangerous diffs do.

The rule of thumb

Proposal at machine speed, approval by an independent mind, human or a different model with an adversarial brief, and never the author merging its own work. The merge right is the ownership right, and the ownership right is the control.

An agent that merges its own pull request has not automated code review. It has automated the removal of code review, and the difference between those two sentences is the incident report.