Reviewing Code the Author Cannot Explain
The review stalled on a strange question. I asked the author what the new error handling does when the downstream times out, and the author, honest and a little embarrassed, said they were not sure, because the block had been accepted from a suggestion and they had not fully read it. The diff was fluent and the tests were green, and neither of those is comprehension. The review I needed to run was not of the code as written but of the comprehension gap around it, and that review has its own checklist.
This is that checklist, written for the increasingly common case where the nominal author cannot narrate the change, and the real author is a model that cannot be asked.
This was a Node 22.14 service, and the change touched a payment path, which is exactly where the gap is most expensive.
1. Interrogate the accepter, not the diff
The first move is to ask the human who merged the change to narrate it, and to treat the narration as the artefact under review. The questions are the same three every time: what does this do at the boundaries, what must it never do, and how did you verify it. If the accepter cannot answer, the review is not failing the code, it is failing the merge, and the correct action is not to review harder but to return the change until someone can own it, per AI wrote it so nobody owns it.
The narration also surfaces the acceptance context: was the suggestion read line by line, or accepted on vibe. The difference predicts the defect density, and it is a fair question, because the accepter is the author.
2. Review the failure paths first, twice
Generated code is disproportionately weak on failure paths, because the training signal and the prompt both centre the happy path. So the review goes straight to the error handling, the timeouts, the retries and the rollback, using the lens from reviewing error handling paths: swallowed errors, retries without budgets, timeouts longer than the caller's. For code the author cannot explain, the failure paths are where the uncomprehended logic bites, and they deserve the majority of the review's attention.
3. Look for the confident unnecessary
Assisted code frequently includes plausible extras: a cache that nothing reads, a lock that guards nothing shared, a config flag that no path checks. These are not malicious, they are the model's prior about what such code looks like, and they are dead weight with failure modes. The review question is, for each non trivial block: which requirement does this serve, and if the accepter cannot name one, it is deleted, because unowned dead code is a trap with a maintenance cost.
4. Demand the invariant, not the implementation
When the author cannot explain how the code works, fall back to what it must guarantee, and review against that. State the invariant in one line, money reconciles, the queue is drained exactly once, the permission is checked on every path, and then verify the code against the invariant by walking the hostile case, not the happy one. Invariant based review works even when the reviewer also does not fully understand the implementation, because it checks the contract, and the contract is the part that matters in the incident.
5. Run the differential
For code the author cannot explain, the cheapest comprehension is execution. Ask for the differential evidence: the test that fails without the change and passes with it, and, for concurrency or timing code, the run against the original untouched test, per the harness lesson in the agent that fixed the bug by weakening the test. If the change cannot produce a failing test for the bug it claims to fix, then nobody, human or model, has demonstrated what it does, and the review has no evidence at all.
6. Check the seams the model cannot see
Models generate against the code they were shown, not against the system that exists. The review must therefore walk the seams the diff touches but cannot see: the callers that now get different behaviour, the queue consumers, the metrics, the migrations, the feature flags. The question is, what outside this file assumed the old behaviour, and the answer usually requires repository knowledge the model lacked and the accepter may lack too, which is the review's most valuable minute.
7. Size the diff to the comprehension
Finally, the review enforces a ratio: the change must be no larger than the accepter can narrate. A thousand line accepted diff that nobody can walk through is not reviewable by definition, and the action is to split it, not to skim it, the same discipline as reviewing large AI pull requests, where the volume is the camouflage. If the team's workflow produces diffs bigger than its comprehension, the workflow is the defect, and the review is the only place that can say so.
The review comment that lands
Not "explain this". Instead: "the timeout on this call is thirty seconds but the caller gives up at five, so the last twenty five are work for nobody; the catch swallows the cause; and no test fails without this change. Until someone can narrate the failure behaviour and produce the failing test, this is not reviewable, it is a claim."
That comment names the gap and the evidence required to close it, and it treats comprehension as the entry fee for merge, which is what it is.
The rule
When the author cannot explain the code, review the comprehension gap directly: interrogate the accepter, double the failure path attention, delete the confident unnecessary, review against the invariant, demand the differential, walk the unseen seams, and split any diff larger than its narration.
The code's fluency is not the risk. The risk is the absence of a person who holds its context, and the review is the last gate where that absence can be priced before it becomes an incident.