The Health Check the CDN Answered Instead of the App

The probe hit the public hostname and the edge returned a cached 200, so the platform believed the service was healthy while every pod behind it was…

Share
The Health Check the CDN Answered Instead of the App. Abstract error autopsy illustration in orange and dark grey on debugly.dev

I have chased this from the incident end and from the down end, and the truth was in the hand-off, as it usually is.

Strip the health check answered by the wrong layer down and one property makes it a false green: the probe asks a public URL, and a cache in front of the application can answer that URL without the application ever being consulted.

This was a Kubernetes service behind a CDN, and every health probe that targets a hostname with a caching layer in front behaves the same way.

Why the probe was answered by the edge

The probe was the HTTP GET, and the GET was the public hostname, and the hostname was the CDN's, and the CDN's was the edge, and the edge was the cache, and the cache was the hit, and the hit was the stored 200, and the stored 200 was the previous healthy response, and the previous was the served, and the served was the probe's answer, and the answer was the not the app's.

The path was the health endpoint, and the endpoint was the cacheable, and the cacheable was the no Cache-Control, and the no Cache-Control was the heuristic, and the heuristic was the edge's guess, and the guess was the TTL, and the TTL was the stored copy, and the copy was the incident, and the incident is the reason the health path must be explicitly uncacheable.

The probe was the correct in form, and the form was the 200's expectation, and the expectation was the met, and the met was the green, and the green was the false, and the false was the layer, and the layer was the invisible, and the invisible is the reason the probe should target the pod rather than the hostname.

Why the green lasted

The cache was the TTL, and the TTL was the minutes, and the minutes was the green's duration, and the duration was the outage's invisibility, and the invisibility was the users' reports, and the reports were the only signal, and the signal was the not the monitoring, and the not monitoring was the incident's length, and the length was the cache's expiry, and the expiry was the red.

The autoscaler read the health, and the health was the green, and the green was the no scale, and the no scale was the no replacement, and the no replacement was the failing pods' persistence, and the persistence was the incident, and the incident is the reason the health signal feeds the decisions, and the decisions are the scaling and the routing, and the two are the wrong when the signal is the cached.

This is the same layering confusion as the cache headers and which one wins, and the common root is the property, not the platform.

The fix, in order

Pointed the probe at the pod, not the public hostname. The probe was the internal address, and the internal was the no cache, and the no cache was the app's answer, and the answer was the truth, and the truth was the fix, and the fix was the target, and the target was the discipline, because the public URL is the layered and the layered is the ambiguous.

Set Cache-Control no-store on the health endpoint. The header was the explicit, and the explicit was the edge's instruction, and the instruction was the no cache, and the no cache was the fix, and the fix was the one header, and the header was the discipline, because the absent header is the heuristic and the heuristic is the guess.

Made the health endpoint return the app's real state. The endpoint was the dependency's check or the process's liveness, and the two were the different questions, and the questions were the readiness versus the liveness, and the versus was the separation, and the separation was the fix, and the fix was the two endpoints.

Verified the probe's path from outside the cluster. The verification was the curl through the CDN, and the through was the comparison, and the comparison was the direct pod's response, and the response was the difference, and the difference was the cache's proof, and the proof was the fix's evidence.

Alerted on the user-facing signal as well as the internal one. The synthetic check was the real user's path, and the path was the cache's inclusion, and the inclusion was the truth, and the truth was the alert, and the alert was the fix, and the fix was the second signal, and the second was the discipline, because the internal probe alone cannot see the edge.

The rule I keep

A health probe that targets a public hostname can be answered by a cache in front of the application, so it reports health the application does not have. Point probes at the pod, set no-store on the health path, separate liveness from readiness, verify the path through the CDN, and keep a synthetic user-facing check as the second signal.

The dashboard stayed green through a full outage because the health probe requested the public URL and the CDN served a cached 200 from before the failure, so nothing restarted, nothing scaled, and the only evidence was the support tickets. That is the whole pattern, and it is the part worth remembering.