A Backup You Have Never Restored Is a Rumour

Share
A Backup You Have Never Restored Is a Rumour. Abstract devops illustration in orange and dark grey on debugly.dev

The incident required a restore, the restore was the plan, and the plan failed in the first hour, on a step nobody had ever executed. The backup system had reported success every night for two years. Both facts are true, and the gap between them is the entire subject: a backup is tested by exactly one operation, the restore, and every night of green checkmarks that never included a restore was a night of untested claims.

This is not an argument that backups lie. It is an argument that the backup and the restore are two different systems, sharing a file format, and only one of them gets exercised.

This was Postgres 16.3 with WAL archiving and nightly base backups, and the drill findings are the standard ones, which is why they belong in one place.

Why the backup succeeds and the restore fails

The backup path and the restore path fail differently, and the backup's success verifies almost none of the restore's prerequisites.

The backup verifies that the data could be read and written out. It does not verify that the archive is complete, because completeness is a property of the restore, which needs every WAL segment between the base backup and the target moment. A pruning job that deleted a segment the restore needs produces a backup that completes nightly and a restore that stops at the gap, and the gap is invisible until the restore walks into it. This is the most common restore failure, and it is silent by construction.

The backup does not verify the restore tooling, the decryption keys, the credentials of the account that reads the archive, or the disk space of the machine that restores, because none of those are touched during backup. The restore fails on a key rotation, an expired credential or a full disk, all of which changed months after the backup began succeeding.

And the backup does not verify the procedure, the order of steps, the configuration of the target, the decisions about which point in time to restore to, because the procedure is executed only during the restore, which is executed only during the incident, which is the worst possible time to discover the procedure is wrong.

The drill is the test

The fix is to execute the restore on a schedule, as a drill, against production backups, into an isolated environment, and to treat the drill's result as the backup's real status. A backup whose drill failed is a failed backup, whatever the nightly checkmark says, and the dashboard should say so, because the dashboard is the only place the truth gets a budget.

The drill need not restore everything every time. A rotation works: the full restore quarterly, the point in time restore to a recent moment monthly, the single table or single database restore weekly, because the partial restore is the more common real need and the cheaper drill. The full restore quarterly is the one that catches the completeness and space questions.

The drill must use the same keys, credentials and tooling the incident would use, from a clean environment, because restoring on the backup server with cached keys verifies less than restoring on a fresh machine the way the incident will. The fresh machine is the point.

What the drill catches, specifically

Run a few drills and the findings arrive in a familiar order.

The pruned WAL segment, which turns the recovery target from a moment into a negotiation, and which the fix addresses by aligning retention with the recovery window the business actually needs, not the window the disk happened to allow.

The key or credential that rotated, which the drill catches in minutes and the incident would catch at two in the morning with the business watching.

The restore time, which is a number the business needs and nobody has. A restore that takes nine hours is a recovery objective of nine hours, and if the business believed it was one, that disagreement is worth discovering in a drill, with coffee, rather than in the incident, without.

The configuration drift in the restore target, the extension, the version, the locale, that makes the restored database subtly wrong, which is the restore equivalent of the helm value you set and the one that applied, where the rebuilt system differs from the original in a way only a comparison catches.

The comparison half

A restore that completes is not a restore that is right. The drill should verify the restored data against the source, a row count, a checksum over a sample, a comparison of a known table, because a restore can succeed structurally and be missing the last hour, which is precisely the hour the incident cares about. The verification step is the difference between "the restore ran" and "the data is back".

The rule

A backup is a hypothesis about a future restore, and the restore is the only experiment that tests it. Drill the restore on a schedule, from a clean environment, with the incident's keys and tooling, verify the data after, and publish the drill's result as the backup's real status.

The nightly checkmark verifies the backup. The drill verifies the plan. The incident will verify whichever one you skipped, and it will do so at the worst time, which is the only time it knows.

The same "the untested path is the one that fails" lesson is the runbook argument in your runbook describes a system you do not have, where the document, like the backup, was complete and unexecuted.

It is also worth naming the organisational half, because the drill dies without it: the restore drill must have an owner and a calendar entry, or it will be the first thing dropped in a busy quarter, precisely the quarter in which the pruning job quietly changes. Pair the drill with the backup's own change feed, so any change to retention, keys, tooling or storage triggers a fresh drill, and the backup's status on the dashboard is the age of the last successful restore, not the age of the last successful backup. One number, the restore's age, carries the whole truth, and a dashboard with it makes the rumour visible every day instead of once a quarter.