The DNS TTL That Was Still an Hour During the Migration
The cutover changed the record and the old address kept serving for an hour, because the TTL was never lowered in advance, and every resolver on the…
Start at the ran and you miss it. Start at the hour of split traffic and you miss how it arrived. The bug lived in the hand-off.
I have hit this enough times that I now check for it before I check for anything clever.
Here is the shape of the TTL that was not lowered in advance, and the one property that makes it unavoidable at cutover time: the TTL is read when the record is cached, not when it is changed, so lowering it after the change has no effect on the copies already held.
The case was a load balancer migration; the mechanics generalise to any DNS cutover.
Why the change did not take effect
The TTL is the cache's lifetime, and the lifetime is the resolver's, and the resolver cached the record an hour before, and the before was the old TTL, and the old TTL was the hour, and the hour was the retention, and the retention was the stale answer, and the stale was the old address, and the old address was the traffic, and the traffic was the split, and the split was the incident.
The TTL change was applied with the record, and the applied was the new TTL, and the new TTL was the five minutes, and the five minutes was the future caches, and the future was the not the existing, and the existing was the hour, and the hour was the already cached, and the cached was the unaffected, and the unaffected was the incident, because the TTL that governs a cached copy is the TTL that was published when the copy was made.
The fix is the advance, and the advance is the TTL lowered before the cutover, and the before is the old TTL's expiry, and the expiry is the five minutes, and the five minutes is the cutover's propagation, and the propagation is the fix, and the fix is the preparation, and the preparation is the day before, and the day before is the discipline.
Why the split traffic was worse than the delay
The split was the two addresses, and the two were the old and the new, and the old was the terminating, and the terminating was the drained, and the drained was the errors, and the errors were the partial, and the partial was the confusing, and the confusing was the diagnosis, because the site worked for some users and not others, and the others was the resolver, and the resolver was the invisible variable.
The split is worse than the outage, because the outage is the clear, and the clear is the fast diagnosis, and the fast is the fix, and the fix is the rollback, and the rollback is the DNS, and the DNS is the same TTL problem, and the problem is the hour, and the hour is the duration, and the duration is the incident's length, and the length is the TTL's fault.
This is the same caching trap as the cached 301 that kept redirecting after the rule was removed, same mechanism, different costume.
The fixes
Lower the TTL a full cycle before the cutover. The TTL drops to five minutes a day ahead, and the day is the old TTL's expiry, and the expiry is the every resolver's refresh, and the refresh is the short cache, and the short is the cutover's speed, and the speed is the fix, and the fix is the calendar entry, and the entry is the discipline.
Keep the old target alive through the window. The old address serves until the propagation completes, and the completes is the TTL's expiry, and the expiry is the safe drain, and the drain is the no split, and the no split is the fix, and the fix is the overlap, and the overlap is the both running, and the both is the discipline.
Verify the propagation from outside. The resolver check is the multiple locations, and the locations are the reality, and the reality is the propagation's state, and the state is the decision, and the decision is the drain's timing, and the timing is the fix, and the fix is the dig from the public resolvers, and the public is the users' view.
Document the TTL in the runbook. The TTL is the cutover's constraint, and the constraint is the plan's input, and the input is the schedule, and the schedule is the window, and the window is the communication, and the communication is the expectation, and the expectation is the fix, because the TTL is the propagation's clock and the clock should be in the plan.
Avoid the CNAME chain that multiplies the TTL. The chain is the multiple records, and the multiple is the multiple caches, and the caches are the longest TTL, and the longest is the propagation, and the propagation is the delay, and the delay is the fix's target, and the target is the flattening, and the flattening is the single record.
The rule
A DNS TTL is read when the record is cached, so lowering it at cutover time does not shorten the copies already held by resolvers. Lower the TTL a full cycle in advance, keep the old target alive through the propagation window, verify from public resolvers, and document the TTL as the cutover's clock.
The record was changed and the traffic kept going to the old address for an hour, because every resolver had cached the record under the old TTL and the new TTL only applied to caches created after the change. Name that one property in review and this class of bug gets hard to ship.