The Rollback That the CDN Kept Serving for an Hour

The deploy was reverted in four minutes, and the incident lasted an hour, because the edge cache held the broken build and kept serving it to everyone the…

Share
The Rollback That the CDN Kept Serving for an Hour. Abstract devops illustration in orange and dark grey on debugly.dev

The rollback was, by the deploy dashboard, a success: the bad commit was reverted, the origin rebuilt, the health checks green, in four minutes. The error rate, however, declined not in four minutes but over the next hour, in a decay curve that matched, exactly, the edge cache's TTL, which is the signature of a rollback that fixed the origin while the edge kept serving the broken build until each cached copy aged out. The incident's duration was a cache configuration value, and nobody had written it down as an incident duration.

This is the anatomy of the rollback that does not roll back for the users, and of the missing control: the purge, which is the rollback's edge half, and which the runbook had never included, because the runbook treated the origin as the system.

This was an origin behind a CDN with an hour TTL on the HTML, and the decay curve is the whole story: the users' experience is the edge's state, and the edge's state is a function of the TTL and the purge, and the purge was absent.

Why the origin fix does not reach the user

The CDN is a cache in front of the origin, and for the cached fraction of requests, which on a well warmed edge is most of them, the origin is never consulted, so the origin's recovery is invisible to those requests until the cached copy expires. The rollback changed the truth at the origin, and the edge kept serving the old truth, and the two truths coexisted for the TTL, with the users' traffic split between them by the cache's luck, which is why the error rate decayed rather than dropped: each expiry returned a user to the fixed origin.

This is the cache as a time machine, serving the past by design, and the design is correct for performance and wrong for the incident, and the resolution is the purge, which is the cache's invalidation, and the invalidation is part of the rollback, not an afterthought.

The rollback's missing half

The runbook's rollback was a deploy action: revert, rebuild, verify origin. The edge was not in the runbook, because the edge is usually invisible, which is its job, and the invisibility is exactly what made it the incident's continuation. The rollback, to be a rollback, must include the edge: purge the affected paths, or the whole edge for a bad build, because a bad build is not a cache to preserve, it is a cache to destroy, and the TTL that protects the origin from load is the same TTL that protects the broken build from the fix.

The purge has its own discipline: a purge of everything is simple and correct for a bad build, and the fear of the purge, that it will cause a cache miss storm and load the origin, is real but mispriced for the incident, because the miss storm hits a fixed origin, which is the healthy state, while the un purged cache hits a broken build, which is the incident, and the trade between a warm broken edge and a cold fixed edge is not close.

The decay curve is the diagnosis

The signature that names this incident is the decay: an error rate that falls over the TTL rather than at the rollback is an edge serving the past, and the curve's half life is the TTL, which is the number to read. This is the measurement discipline from measuring TTFB when a CDN is answering for you, which is why the two fixes look alike.

The same signature distinguishes this from a partial rollback at the origin, which would show a step, not a decay, and the step versus decay read is the partition, as ever, between the origin's state and the edge's.

The fixes that make the rollback whole

The purge is in the runbook. The rollback procedure ends with the edge purge, and the purge is verified by sampling the edge, not the origin, because the edge is the users' truth, and the verification that samples the origin verifies the wrong machine, which is the measurement half of the incident.

The bad build is purged, not expired. A revert of a bad build purges the affected surface immediately, accepting the miss storm against a fixed origin, and the runbook states the trade explicitly, so the operator does not hesitate at the purge prompt during the incident, because hesitation is the TTL, and the TTL is the incident.

The TTL is a reviewed decision. The hour TTL on HTML is a business decision about freshness versus load, and the incident prices it, so the review asks, for each cached surface, how long a broken version may live, which is the TTL's incident meaning, and the surfaces where the answer is minutes get minutes, and the surfaces where the answer is an hour accept the hour, deliberately, with the purge as the escape.

The deploy knows its cache keys. The build tags its output, and the edge can be configured to treat a new build tag as a new object, so a rollback to an old tag serves the old cached objects and a new build never collides, which is versioning the cache, the same discipline as immutable artefacts for the origin, extended to the edge, so the edge's state is a function of the deployed version, not of the clock.

What I would do differently

I would have treated the edge as a deployment target, because it is one: every deploy writes to the edge's state, and a deploy that writes to a store without a rollback story for that store has an incomplete rollback, and the edge is the store, and the store is what the users read, which makes it the only store that matters during the incident.

I would also have graphed the edge's error rate separately from the origin's, so the decay is visible as the divergence of two lines, the fixed origin and the decaying edge, which is the incident drawn rather than inferred, and the drawing is the runbook's first page.

The rule I keep

A rollback fixes the origin, but the users read the edge, and the edge serves the broken build until the TTL or the purge, so the rollback is incomplete without the purge, verified at the edge, and the TTL is the incident duration the runbook never named.

The four minute revert and the hour incident are both true, and the distance between them is the cache, which is the part of the system the deploy writes to and the runbook forgets, and the purge is the one line that makes the two numbers agree.