The AI Suggestion That Rounded the Money and Broke the Reconciliation
The reconciliation job flagged a drift of a few cents across a few hundred orders, and the cents had a pattern, which is how we knew it was not noise. Every affected order had passed through a freshly refactored pricing module, and the refactor had been assisted by an inline code suggestion that a developer accepted with a keystroke. The suggestion was, taken alone, perfectly reasonable. It was also the entire bug.
The module stored money as integer minor units, exactly as it should, per 0.1 plus 0.2 is not 0.3 and your checkout noticed. The accepted suggestion introduced a helper that, for a discount percentage, converted the integer to a float, multiplied, and rounded back. In isolation that helper is fine for many domains. For money, the round trip through a float is the one thing the integer model exists to avoid, and the helper re introduced it in a single line that read like an improvement.
This was Node 22.14, a pricing service, and a reconciliation that compares the charged total to the recomputed total, which is the witness that never lies.
The short answer
The suggestion computed Math.round(units / 100 * percent) / ... style arithmetic through floating point, and for a specific family of values the float round trip lands one cent away from the integer arithmetic the rest of the system uses. The charged amount used the gateway's integer math, the recomputed amount used the new float helper, and the reconciliation saw two truths one cent apart. The defect was not a wrong formula. It was a second, incompatible definition of rounding introduced into a system that had exactly one.
Why every test passed
The unit tests for the module asserted behaviour on round numbers and on the examples in the docstring, none of which sat on a rounding boundary. The property that actually matters, that the helper agrees with the gateway's integer rounding for every minor unit value, was never tested, because it was never stated. This is the coverage gap from code coverage tells you what ran not what was checked: the new line ran, so coverage rose, and nothing verified the invariant it was supposed to preserve.
The suggestion also arrived with the social proof of looking idiomatic. It matched the surrounding style, used a familiar API, and the reviewer's eye slid over it, because reviewing AI assisted diffs already suffers the attention problem in reviewing code the author cannot explain, and a one line "improvement" is the hardest shape to interrogate.
The causes, ranked
1. A second rounding definition
The primary cause. The system had one policy for rounding discounts, integer half up at the minor unit, enforced at the gateway. The helper implemented a different policy via float arithmetic. Two policies for one quantity is a bug by construction, and the float path only agrees with the integer path most of the time, which is the worst possible disagreement profile.
2. No invariant test
There was no test stating that pricing helpers must equal the gateway rounding for all minor unit values in a range. Without that, any refactor can drift the rounding and stay green.
3. Acceptance without narration
The developer accepted the suggestion without being able to state its rounding behaviour under boundaries, which is the ownership gap in AI wrote it so nobody owns it. The keystroke was the merge, and the merge had no reviewer who understood the line.
The fix
The immediate fix deleted the helper and restored the integer path, and the reconciliation cleared within a day. The durable fixes were three.
First, the invariant test: for every minor unit value in a representative range and every discount tier, assert the pricing module and the gateway rounding agree, exactly, as integers. That test now guards every change to the pricing path, and it fails in milliseconds on any future float round trip.
Second, a lint that flags floating point arithmetic in the pricing package, because the package's contract is integers, and a float inside it is a type level violation, not a style choice.
Third, a review rule for accepted suggestions: any accepted suggestion that touches money, auth, or data deletion must be narrated in the pull request, one sentence on what it does at the boundaries, because the sentence is the cheapest possible comprehension check.
What I would do differently
I would have treated the reconciliation as the primary spec for pricing, not as a safety net. The reconciliation already encoded the invariant the system needed, and the bug was invisible everywhere except there. When a system has a component that compares two computations of the same truth, that component is the specification, and tests should be generated from it, not written beside it.
I would also have measured suggestion acceptance the way we measure deploys. A suggestion that changes behaviour is a change, and changes get tests and narration, regardless of whether a human or a model drafted them. The tool that produced the line is not the owner. The person who pressed the key is, and the key press deserves the same rigour as a merge.
The rule
An AI suggestion is a patch from an author who has never read your invariants, and it must be reviewed against the invariants, not against its own plausibility. Money has one rounding policy, and the test that guards it is the spec. Accept the suggestion only when you can state, in one sentence, what it does at the boundary, because that sentence is the difference between a reviewed change and a keystroke.
The reconciliation that caught this is the same discipline as the backup drill in a backup you have never restored is a rumour: the component that checks the truth is the only one that knows the truth.