The Cache Stampede the Second the Key Expired

One popular key expired and four hundred requests missed at the same instant, so all four hundred rebuilt the value, and the database took the whole load at…

Share
The Cache Stampede That Hit the Database the Second the Key Expired. Abstract error autopsy illustration in orange and dark grey on debugly.dev

The spike behaved. The expiry behaved. The chain between them is where working as designed stopped meaning working.

The cache stampede has a single weak property, and here it is: the cache works perfectly until the moment it stops, and the stopping is simultaneous for every waiter, so the load the cache was absorbing arrives at the database in a single instant.

This was a product page cache in front of Postgres, and every cache with a fixed TTL on a hot key behaves the same way.

Why the miss was simultaneous

The key had a TTL, and the TTL was the five minutes, and the five minutes was the expiry, and the expiry was the instant, and the instant was the same for every request, because the requests all read the same key, and the same key was the same expiry, and the same expiry was the simultaneous miss, and the miss was the four hundred, and the four hundred was the stampede.

The traffic was four hundred requests per five minutes on that key, and the traffic was the popularity, and the popularity was the reason the cache existed, and the existed was the protection, and the protection was the expiry's end, and the end was the four hundred at once, and the at once was the database's load, and the load was the query's cost, and the cost was the four hundred times, and the times was the saturation.

The stampede is proportional to the key's popularity, so the hottest key produces the worst stampede, and the worst is the outage, and the outage is the reason the popular keys need a different expiry strategy, because the fixed TTL synchronises the miss, and the synchronised is the spike, and the spike is the incident.

Why the cache made it worse

The cache was the protection, and the protection was the database's capacity, and the capacity was the sized for the misses, and the misses were the rare, and the rare was the sizing, and the sizing was the wrong, because the expiry made the misses not rare, and the not rare was the simultaneous, and the simultaneous was the four hundred, and the four hundred was the capacity's exceed, and the exceed was the saturation, and the saturation was the timeout.

The database was sized for the steady state, and the steady state was the one miss per five minutes, and the one was the cache's rebuild, and the rebuild was the single query, and the single was the fine, and the fine was the assumption, and the assumption was the wrong, because the expiry did not produce one miss, it produced every waiter's miss, and the every was the four hundred, and the four hundred was the incident.

This is the same shape as the Postgres checkpoint storm that looked like a slow disk, and the shared property is the one worth fixing.

The fixes

Locked the rebuild. The first request to miss takes a lock, and the lock is the single rebuild, and the single is the one query, and the others wait for the value, and the wait is the short, and the short is the fix, and the fix is the mutex, and the mutex is the stampede's end, and the end is the prevention.

Served stale while revalidating. The cache keeps the old value past the TTL, and the past is the served, and the served is the no miss, and the no miss is the background rebuild, and the background is the asynchronous, and the asynchronous is the fix, and the fix is the stale-while-revalidate, and the revalidate is the one request, and the one is the database's load.

Jittered the TTL. The expiry has a random offset, and the offset is the spread, and the spread is the desynchronised miss, and the desynchronised is the flat load, and the flat is the fix, and the fix is the plus or minus ten percent, and the percent is the one line, and the line is the prevention.

Warmed the hot keys on a schedule. The popular keys are refreshed before they expire, and the before is the never-missed, and the never-missed is the no stampede, and the no stampede is the fix, and the fix is the cron, and the cron is the schedule, and the schedule is the discipline, because the hot key should never reach its expiry.

Rate limited the rebuild per key. The rebuild is allowed once per window, and the once is the bound, and the bound is the database's protection, and the protection is the fix, and the fix is the limiter, and the limiter is the key's scope, and the scope is the discipline, because the global limiter does not stop the per-key stampede.

The takeaway

A fixed TTL on a hot key expires for every waiter at the same instant, so the whole cached load reaches the database at once. Lock the rebuild, serve stale while revalidating, jitter the TTL, warm the hot keys on a schedule, and rate limit rebuilds per key.

The cache worked for five minutes and then, for one second, did not exist for four hundred requests at the same time, and the database took four hundred identical queries in that second and timed out. I now check for that property first, because it is where the failure actually lives.