Postmortems people actually read
Most postmortems are written for a folder nobody opens. A few structural changes turn them into the most useful document your team produces.
Ask a team where their postmortems live and you’ll usually get a link to a folder. Ask when someone last opened one that wasn’t their own, and the room goes quiet.
That’s a waste. A good postmortem is the cheapest training material you’ll ever get: a real failure, in your real system, with the real sequence of decisions that led there. Teams do write them. They just write them to close a ticket, not to be read.
Write for the engineer who joins next year
The reader you should have in mind is not the incident commander. It’s the engineer who joins in twelve months, gets paged for the same service, and searches for context at 2am.
That reader needs:
- A summary they can read in thirty seconds. What broke, who was affected, for how long, and what fixed it.
- A timeline with decisions, not just events. “14:12: rolled back deploy” is an event. “14:12: rolled back because error rate matched the deploy window; didn’t check the DB yet” is a decision, and it’s the part people learn from.
- What made it hard. Missing dashboards, misleading alerts, unclear ownership. This is where the real findings hide.
Blameless is a structure, not a tone
“Blameless” doesn’t mean being polite. It means the document is structured so that a person is never the root cause.
A useful test: if a sentence reads “X forgot to…”, rewrite it as “the process allowed…”. The deploy checklist didn’t require a migration dry-run is something you can fix. Alex forgot the dry-run is something you can only feel bad about.
People who expect blame hide details. People who don’t, share them. The quality of your postmortems is capped by how safe it is to be honest in them.
Action items that close
The most common failure we see isn’t the write-up. It’s the action items: ten of them, vague, with no owner, parked in a backlog that gets groomed quarterly.
We push teams to a simple rule set:
- Few. Three to five items. If everything is a priority, nothing is.
- Owned. One named owner per item. A person, not a team.
- Dated. A due date, agreed in the review, not assigned later.
- Typed. Tag each item: detect faster, mitigate faster, or prevent. If every item is “prevent”, you’re probably missing cheap wins on detection.
- Tracked in the open. A single view of all open postmortem actions, reviewed in a recurring meeting. Items that slip get discussed, not silently moved.
Make the review worth attending
The review meeting is where a postmortem becomes shared knowledge. Keep it short and structured:
| Block | Time | Purpose |
|---|---|---|
| Summary | 5 min | Everyone gets the same picture. |
| Timeline walk-through | 15 min | Focus on decision points and what people knew at the time. |
| What made it hard | 10 min | Tooling, alerting, process gaps. |
| Actions | 10 min | Agree on owners and dates, live. |
Invite people outside the team that handled the incident. Support engineers in particular often know things about customer impact that never reach the engineering side.
Publish, then point people to it
Finally: make postmortems findable. Link them from the runbook of the affected service. Mention them in on-call handovers. Share a short digest of the month’s incidents with the wider team.
A postmortem that gets read three times has already paid for itself. One that sits in a folder is paperwork, and your team already has enough of that.