A Canary Deployment That Shipped a Broken Release Anyway

Share
A Canary Deployment That Shipped a Broken Release Anyway. Abstract devops illustration in orange and dark grey on debugly.dev

We rolled out a change that broke checkout for a specific class of user. The canary ran for forty five minutes. Every dashboard said healthy. Then we promoted it to a hundred percent and the incident started.

The uncomfortable truth is that the canary worked exactly as configured. The configuration was what failed. A canary that receives too little signal cannot distinguish a healthy release from a broken one, and it will tell you everything is fine with total confidence.

This is a statistics problem wearing a deployment costume.

The arithmetic nobody does

Our service handled about two thousand requests per minute. The canary got five percent, so a hundred requests per minute. Our baseline error rate is around 0.2 percent, and the alert threshold on the canary was a sustained error rate above 2 percent.

Now suppose the bug breaks one request in a hundred, but only for a subset that is ten percent of traffic. In the canary that is one broken request per thousand. Over forty five minutes the canary sees about 4,500 requests and breaks about four or five of them. That is an error rate of 0.1 percent, below even the healthy baseline, and far below the 2 percent threshold.

The canary was never going to fire. The metric could not move that far with that few samples. We had built a smoke detector and installed it in a different building.

This is the same tail blindness that makes average latency meaningless, applied to sampling. The signal exists in the full stream. The canary was looking at a trickle.

Why healthy metrics during a real failure are expected, not surprising

A few structural reasons compound the arithmetic.

Error rate is a ratio. If the bug makes requests slow rather than failing, the error rate does not move at all. A canary watching only 5xx will be blind to a latency regression, and latency regressions are the most common bad release.

The canary traffic is not representative. Routers often send canary traffic that is not a uniform sample. Health checks, synthetic monitors and internal traffic get over represented, and those do not exercise the broken path. Your canary is being tested by the least realistic traffic you have.

Automatic rollback thresholds are tuned to avoid false positives. Teams that were paged by a canary firing on noise raise the threshold. Over months the threshold drifts upward until it only fires on catastrophes. The canary becomes a formality.

What a canary can actually detect

Being honest about the tool, a canary is excellent at detecting large, fast, universal failures. A release that crashes on boot, that returns 500 for everyone, that exhausts memory. Those produce an obvious signal even at five percent, and the canary earns its keep there.

It is poor at detecting small, slow, or segment specific failures. Exactly the kind that cause the most user pain, because they are subtle enough to ship and widespread enough to hurt.

So the design question is not "should we canary". It is "what can this canary see, and what are we pretending it can see".

Making the sample big enough to mean something

There are only a few levers, and they are all about signal per unit time.

Raise the canary share early. Five percent for forty five minutes is a weak sample. Ten to twenty percent for a shorter time is often stronger. The first step is the cheapest place to take risk, because you can roll back fast. A canary that is too small to see is not cautious. It is just blind with extra steps.

Lengthen the observation for slow metrics. If the failure mode is latency or a rare error, the canary needs time to accumulate samples. Match the window to the metric's variance, not to the deploy pipeline's convenience.

Watch absolute counts and segment level metrics, not just the aggregate. If the bug affects a specific route, alert on that route's error count in the canary. A per route or per feature metric has a much higher signal to noise ratio than a global ratio.

Compare the canary against the baseline statistically, not by eye. Two distributions of the same size can be compared with a simple test. If the canary's error count is meaningfully higher than the baseline's over the same window, flag it even below any absolute threshold. This catches the small but real regression that a fixed threshold misses.

The layered defence

A canary is one layer, and it should not be the only one. The release that broke checkout should have been caught by something before it needed the canary.

A feature flag on the new path, so the canary can run the code dark before it affects users. Synthetic transactions that exercise checkout end to end on the canary, so the broken path is hit by traffic designed to hit it. And a fast rollback that makes promotion cheap to reverse, so a mistake costs minutes.

The combination matters more than any single layer. The canary catches the big fast failures. The synthetics catch the specific path failures. The flag catches the ones you can reason about in advance. Together they cover the space that a lone canary only pretends to cover.

The rule

Before you trust a canary, do the arithmetic. Estimate the failure you are worried about, multiply by the canary share and the window, and ask whether the resulting count would move any alert you have. If the answer is no, the canary is a ritual, and you should either change the share, the window or the metric, or admit that rollback speed is your real safety net.

A canary that cannot see the problem is not a safety mechanism. It is a delay with good branding. The same signal blindness shows up in health checks, which is the subject of outage caused by a successful health check.