Blameless Postmortems Went Too Far in the Other Direction

Share
Blameless Postmortems Went Too Far in the Other Direction. Abstract opinion illustration in orange and dark grey on debugly.dev

Here is the thesis. The blameless postmortem movement corrected a real disease, the witch hunt that hid incidents and punished honesty, and then the correction hardened into a culture that will not name a decision, a decider or an owner, and the recurrence rate of the same incident is the price. Blamelessness was a means to get the truth into the room. In many teams it has become a means to keep the consequence out of it, and those are different projects.

I am not arguing for blame. I am arguing that blameless and accountable are not the same word, and that the industry conflated them, and the conflation is costing the same outage, quarterly, with new participants.

What blameless correctly fixed

The old model punished the person nearest the incident, which produced two behaviours that destroyed learning. People hid near misses, because a near miss reported was a risk recorded against a name. And postmortems became prosecutions, which optimise for a verdict rather than an explanation, and a verdict is the least useful output an incident can produce.

Blameless practice, done well, assumes competent people making reasonable decisions with local information, and asks what made the reasonable decision produce the bad outcome. That question is the entire value of the postmortem, and it only gets answered when nobody is on trial. I want that stated plainly, because the argument against the excess is not an argument against the core.

Where it went too far

The excess is the refusal to follow the causal chain into ownership, and it shows in three habits.

The passive voice as a policy. "The deployment was performed" and "the alert was missed" and "the decision was made", with no agent anywhere in the document. The passive voice is not neutrality. It is the removal of the information about who knew what when, which is precisely the information the systemic analysis needs, because decisions are made by people with contexts, and a review that cannot name the context cannot learn from it.

The prohibition on examining judgement. Some cultures made it taboo to ask whether a specific judgement was sound, on the theory that judging the judgement is blame. But judgement is the thing that failed in most incidents, not malice and not stupidity, and a review that cannot examine judgement examines only the machinery, and the machinery is usually fine. The engineer who approved the risky deploy with incomplete information made a bet. The bet is worth understanding, and understanding it is not blaming the bettor.

Action items with no owner, or an owner without authority. The blameless document ends in a list owned by "the team" or by whoever volunteers, and the volunteer has no authority to change the system that caused the incident, so the list decays. Accountability is not punishment. It is the assignment of the consequence to a person with the power to act, and its absence is not kindness, it is abdication.

The recurrence is the evidence

The test of a postmortem culture is the recurrence rate of the same class of incident, and the over corrected culture fails it visibly. The incident whose root cause is a decision nobody can name, repeated quarterly with a new author each time, is the signature. Each review is honest, thorough and blameless, and each produces recommendations, and the decision recurs, because the review never reached the point where a person with authority owns the change, and ownership is the only mechanism that converts a lesson into a different decision next time.

This is the same loop as the people who designed it should be on the pager, where the feedback is collected and then routed away from the person who can act. The blameless excess is that routing, formalised into the review itself.

The counterargument, fairly

The defence of the strict reading is that any examination of individual judgement will, in practice, slide back into punishment, because organisations are not trustworthy with the distinction, and the safest equilibrium is the blanket taboo. I take that seriously. The taboo is a load bearing wall in organisations with a punitive history, and removing it carelessly rebuilds the witch hunt.

But a wall that was built to keep out the prosecutor should not also keep out the accountant. The mature practice distinguishes them: no punishment for honest error, and full, named examination of judgement, and an owner with authority for every consequence. The teams that hold that line get both the truth in the room and the change after it, and their recurrence rates show it.

What accountability without blame looks like

Concretely, the review names the decisions and their authors, in neutral language, because the name is data. It asks, for the key decision, what the decider knew, what they expected, and what would have changed their mind, which is the systemic question with the person in it. It ends with action items each owned by one person with the authority to complete them, and the owner reports completion, and the completion is checked against the recurrence, not against the document.

And it keeps the guarantee that makes honesty safe: an honest error, reported, is never punished. The guarantee is the price of the truth, and it is non negotiable. What is negotiable, and negotiated away in the excess, is the second half, that the truth must lead to an owned change.

The rule of thumb

Blameless answers the question "what should we learn" and accountability answers "who makes the change", and a postmortem that asks only the first produces a document, while one that asks both produces a different next quarter.

Keep the guarantee that protects honesty. Restore the examination of judgement that protects learning. And treat the recurrence of the same incident as the audit of the review culture, because it is the only audit that cannot be written in the passive voice.

The first minutes of the incident are where the culture shows itself live, in the first ten minutes of an incident, and the dashboard version of the same abdication is nobody reads your dashboards.