The Checkpoint Storm That Spiked Latency Every Few Minutes
Latency was fine, then terrible, then fine, on a rhythm of a few minutes, with no traffic rhythm to match. The period was the checkpoint interval, and the…
The p99 chart had a heartbeat: a spike every few minutes, sharp, regular, unrelated to traffic, which is the signature of a background process with a schedule rather than a workload problem, because workloads do not pulse on a clock. The period matched the checkpoint interval, and the spikes were checkpoints, the database periodically flushing every dirty page it had accumulated, and the flush, configured too aggressively and too fast, was competing with foreground queries for the same disks, which is how a housekeeping task becomes a latency incident.
I have lost on-call weeks to this exact shape of slow burn, so the order below is the order I fix it in.
This is the mechanics of the checkpoint, and of the storm, because the checkpoint is one of the few background processes that can directly tax the foreground, and the configuration that controls the trade is rarely touched until it hurts.
This was Postgres 16.3 on spinning behaviour emulated by throttled disks, a write heavy workload, and the spikes arrived every five minutes, which is the default checkpoint period, which is the first clue that a default was doing the deciding.
What a checkpoint is for
Postgres writes changes to the WAL first, which makes commits fast and durable, and applies the changes to the heap pages in shared buffers lazily, which leaves dirty pages in memory. The checkpoint is the moment those dirty pages are flushed to the heap on disk, and it is also the moment the WAL becomes recyclable, because everything before the checkpoint is guaranteed persisted. So the checkpoint is the bridge between the fast, sequential WAL and the slow, random heap, and it must happen, periodically, or the WAL grows forever and recovery after a crash replays forever.
The checkpoint is therefore not optional. The question is its cost, and the cost is that flushing thousands of dirty pages is a burst of random writes, and if the burst lands while foreground queries need the same disks, the queries wait, which is the spike.
Why the storm forms
The storm is the configuration failing to spread the work, and it forms from three settings acting together.
The interval. A short checkpoint_timeout means checkpoints often, so each accumulates fewer dirty pages but the bursts arrive frequently, and the frequent bursts mean the disks are repeatedly stolen, producing the heartbeat in the latency chart. The five minute default with a write heavy load is a frequent burst.
The completion target. The spread of the flush is governed by how fast Postgres is allowed to write during the checkpoint, historically the checkpoint completion target. If the flush is allowed to run fast, it finishes quickly but hogs the disks while it runs, a tall spike. If spread slowly, the spike flattens but the flush runs longer, risking it not finishing before the next checkpoint, which is the next failure mode.
The volume of dirty pages. A write heavy workload with too little shared buffers accumulates a large dirty set per interval, so each checkpoint is large, and the large flush is the tall spike. The size of the spike is the workload, the frequency is the interval, and the shape is the completion target, and the three multiply.
Reading the storm
The confirmation is in the checkpoint statistics. Postgres logs checkpoint activity when enabled, reporting the number of buffers written and the time, and a checkpoint that writes a large fraction of shared buffers in a few seconds is the spike. The metric to graph is buffers written per checkpoint and the checkpoint write time, against the foreground p99, and the correlation, spike aligned with checkpoint, is the diagnosis, because no other background task has that period.
The distinction from autovacuum is worth keeping: autovacuum's bloat is a slow creep, per autovacuum is running and your table is still bloating, while the checkpoint is a periodic spike, and the two are the housekeeping pair, one pruning dead tuples, one flushing dirty pages, and both can tax the foreground when misconfigured.
The fixes
Spread the flush. Raise the completion target's spread so the checkpoint writes slowly across the interval, trading a tall short spike for a low long one, which is the foreground friendly shape. The flush must still finish before the next checkpoint, so the spread is bounded by the interval, and the two are tuned together.
Lengthen the interval and enlarge the buffers. A longer checkpoint_timeout accumulates more dirty pages per checkpoint but flushes less often, and larger shared buffers hold the dirty set comfortably, so the fewer flushes are absorbed. The product, interval times spread, must exceed the flush size, which is the capacity arithmetic of housekeeping.
Align the disks. The storm is worst when the WAL and the heap share slow disks, because the checkpoint writes to the heap while commits write to the WAL, and the two compete. Separating them, or giving the heap fast storage, removes the contention that turns a flush into a foreground stall.
Watch the checkpoint's share. Graph the fraction of buffers written by checkpoints versus the background writer, because a healthy system spreads most writes through the background writer continuously, leaving the checkpoint with little to do, and a storm is the inversion, the checkpoint doing the bulk of the writing in bursts, which is the shape to alert on.
The takeaway
A checkpoint is the necessary bridge between the fast WAL and the slow heap, and its flush is a burst of random writes that, unspread, lands in the foreground's path on a clock, producing the periodic spike no traffic chart explains. Spread the flush, lengthen the interval with larger buffers, separate the WAL from the heap, and alert on the checkpoint's share of the writing.
The heartbeat in the p99 was not the workload. It was the housekeeping, and the housekeeping had a schedule, and the schedule, unchecked, is the storm, which is the oldest lesson in this archive: the background is part of the system, and the system includes your p99.