The 301 That Cached Itself and Became a Loop

A misconfiguration emitted a permanent redirect from a URL to the same URL for twenty minutes, and the browser cached the 301, so the loop kept spinning for…

Share
The 301 That Cached Itself and Became a Loop. Abstract bug hunt illustration in orange and dark grey on debugly.dev

The fix deployed at noon and the incident continued until the next morning, because the loop was no longer on the server. For twenty minutes a misconfiguration had answered requests to a path with a 301 to the same path, a perfect loop, and the browsers that visited during those twenty minutes cached the redirect, as browsers are told to do for permanent redirects, and kept honouring it after the server forgot it ever happened. The server was fixed. The clients were not, and the clients are the ones that spin.

This is the bug hunt for a redirect that outlived its cause, and for the specific property of the 301, permanence, that turns a brief misconfiguration into a durable client side loop.

This was an edge config that mapped a path to itself during a botched rewrite, observed in Chrome 133 and Safari 17, and the loop's afterlife is the caching half of the story, which is the half the server team cannot see.

The symptom, precisely

After the fix, new visitors and fresh profiles loaded fine, and returning visitors, and the same profiles, spun in the loop, which is the selectivity that names a client cache: the server answers correctly for everyone, but a subset of clients never asks, because their cache answers first with the stored 301, which points at the same URL, which the cache answers again, and the browser, following its own cache, loops without a single request reaching the fixed server, which is why the server logs showed nothing and the users showed everything.

The hypotheses that died

The fix did not actually fix

First I rechecked the server, curling the path from a clean client, which returned 200. It died because the server was fine for anyone without the cache.

A CDN still served the redirect

Next I purged and rechecked the edge, which also returned 200. It died because the loop persisted in browsers with the edge clean, so the cache was below the edge, in the browser.

A service worker or extension

Then I suspected a worker rewriting. Disabling extensions and workers changed nothing for the affected profiles. It died because the cache involved was the ordinary HTTP cache, the most boring one, doing its specified job.

The breakthrough

The diagnosis is the status code's semantics. A 301 is permanent, and browsers treat permanence as a long lived cache entry for the redirect itself, so the mapping from URL to target is stored and reused without revalidation, and a 301 from a URL to itself, cached, is a fixed point the browser will honour indefinitely, or until the cache is cleared, which the user must do, because the site cannot reach into the client's cache to apologise. The twenty minute misconfiguration was thus amplified by the permanence semantics into a per client incident with no server side remedy, which is the asymmetry that makes this bug nasty: the cause is server side and brief, the effect is client side and long, and the two are decoupled by the cache.

This is the cache as accomplice pattern from you fixed CORS and the browser kept failing for another ten minutes, where the client's cache extends the incident past the fix, but with a 301 the extension is unbounded rather than ten minutes, because permanent means permanent.

Why the loop formed in the first place

The server side half is also worth the autopsy: a rewrite rule that mapped a path to itself, which is a one line config error, but the loop it created was only briefly harmful server side, because a server side loop of 301s is bounded by the browser's redirect limit, which stops the spin after a handful of hops and shows an error. The durable harm came from caching, not from the loop, and the config error's severity was therefore mispriced at review: a self redirect looks like a harmless typo, and it is a typo that plants a fixed point in every visitor's cache, which is a typo with a blast radius.

What actually fixed it

The immediate client side remedy is not yours. The affected users needed a cache clear, or the cache entry's expiry, and the honest fix for the population is time, which is an unsatisfying remedy and the reason the prevention matters. Some teams ship a redirect that breaks the loop, a 301 from the URL to a different URL then back, to overwrite the cached fixed point, which works where you can reach the clients, and is a confession that the cache is the incident.

Never emit a 301 to the same URL, and validate rewrites against self loops. The config layer should reject a rule whose target equals its source, because the self loop is the fixed point, and the check is one comparison at deploy, which is the cheapest prevention in this post.

Use 302 for anything you are not certain is permanent. The permanence semantics are what cached the loop, and a 302, temporary, is cached far less aggressively, so the misconfiguration's afterlife shrinks from indefinite to brief, and the default should be temporary unless permanence is a decision, because permanence is the property that turned twenty minutes into a day.

Alert on redirect loops at the edge. A loop is detectable server side while it exists, a request whose redirect chain revisits a URL, and alerting on it catches the misconfiguration during its brief server side life, before the cache distributes it, which is the tripwire placed at the only moment the server still has agency.

What I would do differently

I would have priced the status code as a cache instruction, not as a semantics nicety, because 301 versus 302 is a decision about how long a mistake lives in a client, and decisions about mistake lifetime belong in the review, and the review that approved the rewrite never asked how long its failure would persist, which is the question the incident answered: a day, in every browser that visited.

I would also have made the cache's afterlife visible in the incident tooling, a way to know how many clients hold the entry, which is unanswerable from the server, and the unanswerability is the lesson: once you hand a decision to the client's cache, you lose the ability to observe, and then to revoke, that decision.

The takeaway

A 301 is a permanent instruction that clients cache without revalidation, so a brief self redirect plants a fixed point in every visitor's cache that outlives the server fix indefinitely. Reject self loops at deploy, default to 302 unless permanence is deliberate, and alert on redirect chains that revisit a URL while the server still has agency.

The loop was fixed at noon and raged until morning, and the distance between those two facts is the client's cache, which is the part of your system you deploy to but cannot patch, which is why the permanence of what you hand it deserves the review the incident never got.