The Shopify Webhook Retry That Double Booked the Order
The delivery timed out after the handler had already run, so the platform retried, and the handler, which assumed delivery meant not processed, booked the…
The inventory showed two bookings for one order, and the order history showed one, which is the signature of a side effect executed twice with the record written once, and the cause was a retry the handler never expected: the webhook delivery had timed out at the platform's deadline, but the handler had already done the work and was merely slow to answer, so the platform, seeing no acknowledgement, did what delivery systems do, and redelivered, and the handler, treating each delivery as a new event, booked again.
This is the anatomy of the duplicate delivery, the at least once contract meeting a not idempotent handler, and of the idempotency key that is the only durable fix.
This was a Shopify orders webhook consumed by a Node 22.14 handler that adjusted inventory, and the double booking is the platform's retry semantics working as documented against a handler that assumed a semantics it was never given.
The contract is at least once, and that is the platform's right
A webhook delivery system cannot distinguish a handler that failed from a handler that succeeded and lost the acknowledgement, because both look like silence, so it must redeliver on timeout, and it does, with retries over a schedule, because the alternative, never retrying, loses events, which is the worse failure, per shopify webhooks not received debugging, the missing delivery cousin of this duplicate one. The contract is therefore at least once, and the handler's obligation, which the platform states and the incident proved, is to make the second delivery a no op, which is idempotency, and it is the consumer's job, always, because the producer's job is to not lose.
Why the handler doubled the work
The handler read the payload and applied the change, with no memory of having seen the event before, because it was written against the mental model of exactly once, which is the model every developer assumes and no delivery system provides. The doubling was not a bug in the booking logic, which was correct per call, it was the absence of a deduplication step, and the absence is invisible until the first retry, which arrives, per the schedule, minutes after the first, which is why the double booking appeared as a surprise rather than at launch, because launches rarely hit the timeout that triggers the retry.
This is the same shape as the webhook idempotency and duplicate delivery post's core, and the ecommerce variant adds the money dimension: the duplicated side effect is inventory and, in the payment webhook cousin, a charge, so the duplicate is not a log line, it is a reconciliation incident.
The fixes that make the retry a no op
The idempotency key is the event identity. Each delivery carries an event id, and the handler records processed ids, in a table or a store with a unique constraint, and the booking happens only if the insert of the id succeeds, so the second delivery's insert violates the unique constraint and the handler returns success without booking, which is the deduplication made structural, the unique constraint doing the work the memory could not, per the ledger discipline in the agent ran the migration twice because it misread its own log, where the ledger is what turns a retry into a no op.
Acknowledge fast, work async. The timeout that triggered the retry was the handler doing the work synchronously before answering, so the fix is to record the event and acknowledge immediately, then perform the booking on a queue, so the delivery's acknowledgement is fast and the retry, if it comes, finds the id recorded and no-ops, separating the delivery contract from the work's latency, which removes the timeout as a source of duplicates at the root.
Make the side effect itself safe. Where the booking can be expressed as an upsert or a delta keyed on the order, the repeat applies the same delta to the same key and the second application is detectable or harmless, defence in depth beneath the idempotency key, because the key can be bypassed by a bug and the structural safety remains.
The test that catches it
The regression test delivers the same event twice, with a timeout simulated between the work and the ack, and asserts the inventory moved once, and the assertion on the second delivery's no op is the contract test, and it belongs in the suite the day the handler is written, because the retry is not an edge case, it is the platform's documented behaviour, and a handler tested only against single deliveries is tested against a world the platform does not provide.
It is also worth logging the duplicate rate, deliveries that no-opped against the id table, because the rate is the platform's retry traffic made visible, and a spike in duplicates is a latency problem in the handler, the ack slowing, which is the early warning that the next duplicate will find a gap.
The rule I keep
Webhook delivery is at least once by the producer's necessity, so the consumer must make the second delivery a no op, with an idempotency key enforced by a unique constraint, an acknowledgement that precedes the work, and a side effect expressed so a repeat is harmless.
The retry was never the bug. The retry is the delivery system keeping its promise not to lose, and the double booking was the handler breaking its promise not to repeat, and the distance between those two promises is exactly the idempotency key, which is the shortest contract in distributed systems and the most skipped.
It is worth naming the reconciliation half, because the duplicate that slips past is found by the money, not the logs. A nightly join of side effects against event ids, flagging any order whose booking count exceeds one, is the audit that catches the gap in the id table before the customer does, and the audit's existence is what makes the idempotency key a defended boundary rather than a hopeful one, turning the rare duplicate from an incident into a row in a report that someone reads.