> ## Content Index
> Fetch the complete content index at: https://debugly.dev/llms.txt
> Use this file to discover other available public pages before exploring further.

# Why Adding One Server Moved Everyone's Cache
- URL: https://debugly.dev/one-server-moved-every-cache/
- Published: 2026-10-09T11:57:00.000Z
- Updated: 2026-10-10T14:00:02.000Z
- Description: The cache was sharded by modulo, so the twelfth server rehashed every key, and the whole cluster went cold at once, and the database took the entire load in…
- Author: Rohit Bhadani
- Tags: Deep Dive, Caching, Distributed Systems

I would like to say the modulo was the problem, because that would be a shorter post, but it started at the scaling.

Strip the modulo sharding rebalance down and one property makes it catastrophic: the shard is the hash modulo the server count, so changing the count changes the shard for almost every key, and the whole cache is invalidated by the change that was supposed to help.

This applies to any cache or sharded store keyed by hash modulo N, and the mechanism is the arithmetic, which is the same regardless of the implementation.

## Why the modulo rehashes everything

The shard is the hash of the key modulo the count, and the count was the eleven, and the eleven was the divisor, and the divisor was the shard's determinant, and the determinant was the twelfth's change, and the change was the divisor's, and the divisor's was the every result, and the every result was the different shard, and the different shard was the miss.

The fraction that stays is the one over the new count, and the new count was the twelve, and the twelve was the one twelfth, and the one twelfth was the eight percent, and the eight percent was the hit rate, and the hit rate was the collapse, and the collapse was the ninety two percent's miss, and the miss was the database, and the database was the incident.

The arithmetic is the reason, and the reason is the not the implementation's bug, and the not bug is the design, and the design is the fix's target, and the target is the consistent hashing, and the consistent hashing is the minimal movement, and the movement is the one over the count, and the count is the fix.

## Why the database could not absorb it

The cache was the database's protection, and the protection was the ninety five percent, and the ninety five percent was the absorbed, and the absorbed was the database's sizing, and the sizing was the five percent, and the five percent was the capacity, and the capacity was the not the hundred, and the not hundred was the saturation, and the saturation was the timeout, and the timeout was the incident.

The stampede was the simultaneous, and the simultaneous was the every key's miss at the same instant, and the instant was the rebalance, and the rebalance was the no stagger, and the no stagger was the full load, and the full load was the collapse, and the collapse is the reason the rebalance should be gradual, and the gradual is the fix.

This is the same shape as [the cache stampede that hit the database the second the key expired](https://debugly.dev/cache-stampede-after-key-expiry/), and the shared property is the one worth fixing.

## What actually fixed it

**Used consistent hashing.** The ring was the mapping, and the mapping was the minimal movement, and the movement was the one over the count, and the count was the twelve's twelfth, and the twelfth was the eight percent's invalidation, and the invalidation was the survivable, and the survivable was the fix, and the fix was the algorithm.

**Used the bounded-load variant.** The bound was the hot shard's overflow, and the overflow was the next node, and the next node was the balance, and the balance was the fix, and the fix was the discipline, because the plain ring puts the hot key's whole load on the one node.

**Warmed the new node before the traffic moved.** The warm was the preloaded, and the preloaded was the miss's absence, and the absence was the database's protection, and the protection was the fix, and the fix was the sequence, and the sequence was the discipline, because the cold node is the stampede's source.

**Moved the traffic gradually.** The gradual was the percentage, and the percentage was the partial rehash, and the partial was the controlled miss rate, and the controlled was the fix, and the fix was the rollout, and the rollout was the discipline, because the instant switch is the instant full miss.

**Kept the old shard readable during the migration.** The dual read was the new's miss then the old's lookup, and the lookup was the hit, and the hit was the database's protection, and the protection was the fix, and the fix was the transition, and the transition was the discipline, because the cold cache should read the old location rather than the database.

## The rule

Sharding by hash modulo the node count rehashes almost every key when the count changes, so scaling the cache invalidates it. Use consistent hashing with bounded loads, warm the new node before moving traffic, migrate gradually, and read the old shard during the transition.

Adding the twelfth cache server dropped the hit rate from ninety-five percent to about eight percent in one second and saturated the database, because the shard was the hash modulo the server count and changing the count changed the shard for eleven keys in twelve. Name that one property in review and this class of bug gets hard to ship.