A postmortem can become a useful engineering document or a very expensive way to write “we should improve monitoring”. The difference appears after the meeting: someone changes a specific part of the system, tests that change and checks whether the operational risk actually decreased.
This guide uses a fictional checkout incident to show the whole path, from collecting evidence to verifying follow-up work. The example names, numbers and dates are invented for illustration; they do not describe an IncidentRelay outage or a customer incident.
A good review should help a responder who was not there understand what customers experienced, why the response unfolded as it did and which safeguards will behave differently next time.
Decide what deserves a review before the next outage
Agree on lightweight triggers while production is behaving. Examples include material customer impact, an unexpectedly difficult recovery, a failure of detection, or a near miss that exposed a weak control. Leave room for responders to request a review even when an incident misses a numerical threshold.
Google's Postmortem Culture chapter recommends defining review criteria in advance and investigating contributing conditions without blaming individuals. The practical benefit is better evidence: people need to be able to describe uncertainty and mistakes without turning the review into a defense of their reputation.
| Situation | Useful starting scope | What to preserve |
|---|---|---|
| A short, understood interruption | A concise written review. | Impact, evidence, recovery and a decision about follow-up. |
| A significant outage or difficult recovery | A written investigation and facilitated discussion. | Decision context, competing explanations and agreed improvements. |
| A near miss or unexpected successful recovery | A focused learning review. | The safeguard that worked, its limits and whether it can be repeated. |
| A recurring failure already reviewed | A linked follow-up to the earlier review. | Why prior actions were insufficient, unfinished or ineffective. |
Scale effort to what remains uncertain and what can be learned. A ten-minute incident can expose an important systemic weakness. A long review does not automatically provide a better explanation; sometimes it provides the same explanation in a larger font collection.
Capture evidence while the response is still fresh
After stabilizing the service, assign a review owner and collect the material most likely to disappear: relevant log windows, configuration revisions, deployment identifiers, alert history, responder notes and the queries behind impact estimates. Let exhausted responders recover before expecting them to reconstruct every decision.
Record a source and a time basis for each important statement. “The dashboard was red” is less useful than a saved query for a named service, region and interval. Keep the time range and filters with a chart; tomorrow's default dashboard view may no longer show yesterday's incident.
Use one timezone in the timeline, preferably with explicit UTC offsets. Separate the time an event happened from the time someone observed or recorded it. If clocks disagree, document the discrepancy instead of rearranging timestamps until the story looks tidy.
Preserve only the evidence needed for the investigation in an appropriate location. Redact credentials and unnecessary customer identifiers from shared extracts, and check that the intended reviewers can open the remaining links.
A worked example: a checkout configuration change
In this fictional incident, a checkout service deploys a database connection-pool limit change on 22 September 2026. The change reduces the per-process limit from 80 to 20. It passed a low-concurrency staging check, but production traffic spends longer holding connections than the test workload did.
From 14:02 to 14:15 UTC, the affected regional checkout endpoint records 12,000 request attempts, including retries. Of those attempts, 2,400 return a timeout response: 20% of attempts in that interval. These numbers do not establish how many distinct customers were affected or how many purchases were ultimately abandoned.
A separate asynchronous queue returns to its normal age at 14:27. That is a later recovery milestone than the end of elevated HTTP errors. Keep both visible rather than choosing whichever timestamp produces the nicer duration.
| Time (UTC) | Observation or action | Evidence and interpretation |
|---|---|---|
| 14:00 | The pool-limit configuration rolls out. | Deployment record and configuration diff establish the change. |
| 14:02 | Checkout timeout responses and connection-acquisition waits rise. | Service metrics establish the first observed impact; causality still needs investigation. |
| 14:04 | The checkout alert fires and a notification is delivered. | Monitoring and delivery records establish detection and paging separately. |
| 14:06 | The on-call responder acknowledges and checks database health. | The incident event and a responder note establish ownership and the initial hypothesis. |
| 14:10 | The team begins reverting the pool-limit change. | The response log records the decision and rollback job. |
| 14:14 | The previous configuration is active across the affected deployment. | Runtime configuration checks confirm the rollback completed. |
| 14:15 | Elevated checkout timeout responses stop. | The affected endpoint's error series returns to its baseline. |
| 14:20 | The team confirms five minutes of stable checkout behavior. | A follow-up check supports continuing recovery; it does not erase the queue backlog. |
| 14:27 | The asynchronous queue age returns to normal. | Queue telemetry establishes the second recovery milestone. |
The timeline is an index into evidence, not a transcript of every chat message. Include decisions that changed the response, hypotheses that consumed meaningful time and observations that altered the team's understanding.
Write impact before explaining the cause
Describe the affected user journey, scope and measurement. In the example, “checkout timeout responses in one region” is more precise than “the platform was down”. Report the denominator with every percentage and say whether it counts attempts, unique operations, accounts or another unit.
Record what you have not established. The fictional response logs demonstrate failed attempts, but they do not by themselves prove customer abandonment or the absence of duplicate orders. Assign reconciliation work if those questions matter, and avoid replacing missing evidence with “no data loss”.
Use several labeled intervals where needed: first observed customer impact, alert detection, acknowledgement, mitigation, stable service and backlog recovery. An alert's resolved timestamp is evidence about the alert lifecycle; customer recovery may require additional checks.
Keep facts, inferences and unknowns visibly different
| Type | Example statement | Next step |
|---|---|---|
| Confirmed fact | The deployed configuration limited the pool to 20, and acquisition waits rose after rollout. | Link the configuration revision and time-series query. |
| Supported explanation | A later isolated replay reproduces the wait increase with limit 20 and removes it with the prior limit under the same workload. | Save the replay setup, observations and limitations. |
| Unquantified hypothesis | Client retries may have increased the load while checkout was slow. | Compare attempt identifiers and retry telemetry before estimating the contribution. |
| Open question | The number of abandoned purchases is not yet known. | Name an owner and deadline for the impact reconciliation. |
Changing your mind during review is progress when the evidence improves. Record the correction and why it happened. Do not preserve a confident early theory merely because it already appears in three incident updates.
Explain the conditions that made the incident possible
The trigger in the example is the configuration rollout. The failure mechanism is insufficient connection-pool capacity for the production workload, producing waits and request timeouts. Contributing conditions explain why the change reached users and why recovery took the time it did.
Those conditions might include a staging workload with too little concurrency, a rollout check that observes process health without exercising checkout, or an operational dashboard that makes connection-acquisition waits hard to find. Include each only when investigation supports it.
Ask what the engineer and reviewer could see at the time: which tests passed, what problem the change was intended to solve, which assumptions were reasonable and which information was unavailable. The relevant design question is how the next engineer could make a better decision under similar conditions.
“The engineer chose the wrong value” ends the explanation at the point where useful engineering questions begin.
A stronger account is: “The approval check did not exercise the production concurrency range, so the lower limit passed without revealing the connection wait.” That statement points toward a testable safeguard and can be challenged with evidence.
Do not force every incident into one root-cause box. The initiating change, the conditions that amplified it, the detection gap and the recovery friction can each need different work. Google's published postmortem analysis separates triggers from contributing cause categories; that distinction is useful for recognizing recurring patterns across reviews.
Run a review that improves the written account
Circulate the draft before the meeting. Ask participants to mark factual corrections, missing decision context and unsupported claims. A facilitator can keep the discussion focused; the review owner records decisions and maintains the document. These responsibilities do not have to fall to the person who deployed the change.
For a small incident, a 30- to 45-minute meeting is a reasonable starting point:
- Confirm the scope and impact, including the measurements that remain uncertain.
- Walk through the decision points and correct the timeline.
- Discuss the failure mechanism, contributing conditions and safeguards that helped.
- Select a small set of follow-up actions and agree on owners and verification.
- Read back open questions and the next review date.
When a discussion needs an experiment, record the experiment instead of conducting a twenty-minute contest of increasingly confident guesses. A blameless review still needs clear accountability for investigation and follow-up.
Ticket: “Improve monitoring.”
Future on-call: “Delighted to inherit this measurable unit of hope.”
Give every action an owner and a proof of completion
Write actions around the failure mode they address. Choose work that prevents an unsafe change, detects customer impact sooner or makes recovery more reliable. Assign one accountable owner, a due date, a reviewer and a concrete acceptance check. A team can help; “the team” is not a person who can explain a missed deadline.
Here is an illustrative follow-up plan for the fictional incident. The thresholds are example acceptance criteria for this service, not general defaults for every checkout system.
| Action and purpose | Owner and due date | Acceptance evidence and reviewer |
|---|---|---|
| Add a representative concurrency check to pool-configuration changes, preventing the same unsafe approval. | Mira, 2 October 2026. | The saved workload rejects the incident configuration and passes the agreed safe baseline; under that baseline, timeout rate stays below 1% and acquisition-wait p95 below 100 ms. Chen reviews the run and release gate. |
| Validate the checkout impact alert and its delivery path, reducing time to a useful page. | Owen, 30 September 2026. | A controlled test that breaches the rule's sustained-impact condition reaches the expected responder within two minutes; a healthy control produces no page. Mira reviews timestamps and routing. |
| Exercise the configuration rollback procedure, making recovery repeatable. | Chen, 6 October 2026. | A responder who did not write the procedure restores the expected configuration hash within five minutes in a safe exercise and checks checkout health and queue recovery. Owen records the result. |
Before accepting an action, check that its owner has capacity and authority to complete it. Prioritize by expected reduction in impact, likelihood or response difficulty. If an action is deferred, record the residual risk, the person accepting the decision and a date to revisit it.
“Add a test” is incomplete until you know what would fail. “Update the runbook” is incomplete until someone can follow it. A merged change can be a step toward completion; the verification result establishes whether the intended control works.
Use IncidentRelay as a source of response evidence
In IncidentRelay 2.2.0, alert-group comments can hold investigation notes, decisions and links to logs, tickets or a review document. System-generated alert events record lifecycle activity. Use both: a timestamped ACK shows when ownership was recorded, while a comment can explain what the responder believed and chose to investigate.
Explain Trace can help answer why an incoming alert was routed, grouped, suppressed or stopped. It is useful when the response problem involved paging or routing, but it does not establish the root cause of a database wait or a checkout timeout. Pair it with the relevant service telemetry.
Keep the reviewed postmortem in the document or issue system your team already maintains, and put its location in the incident notes. Check evidence retention before relying on incident history as a permanent archive: configured cleanup of resolved alert groups also removes dependent records such as comments and lifecycle events.
These workflow details were checked against the published v2.2.0 source. The comments guide, Explain Trace guide and retention guide describe the relevant behavior and settings.
A postmortem template you can copy
Use this as a starting document. Replace the placeholders with evidence and decisions; remove sections that do not help the review. Keep an unresolved question explicit rather than filling every field with a guess.
# Incident review: [service and user-visible failure]
Review owner: [person]
Status: [draft / reviewed / follow-up open / verified]
Incident reference: [link or identifier]
Incident date: [YYYY-MM-DD]
Review date: [YYYY-MM-DD]
Timeline timezone: [UTC or explicit offset]
Reviewers: [people and relevant roles]
## What users experienced
Affected journey, service and scope: [...]
First observed impact: [...]
End of elevated errors / stable service: [...]
Other recovery milestones: [...]
Measurement: [count, denominator, interval, query]
Known exclusions and uncertainty: [...]
## Summary
[What happened, how impact was reduced, current state.]
## Evidence register
E1: [source, exact interval/filter/revision, location]
E2: [...]
[Note gaps, retention limits and access restrictions.]
## Timeline
[Time | observation/action | evidence | decision context]
[Distinguish event time from observation time.]
## Explanation
Trigger: [...]
Failure mechanism: [...]
Contributing conditions and evidence: [...]
Why existing safeguards did not catch or contain it: [...]
Alternative explanations and how they were tested: [...]
## Response
What helped detection, coordination and recovery: [...]
Where responders lacked information or a safe action: [...]
What assumptions were reasonable at the time: [...]
## Open questions
[Question | investigation owner | due date | evidence needed]
## Follow-up actions
[Ticket | risk addressed | action | owner | due date]
[Acceptance test | reviewer | verification evidence]
[Priority and dependencies; deferrals with a review date.]
## Review and follow-through
Factual corrections agreed: [...]
Remaining disagreements or uncertainty: [...]
Next action review date: [...]
Verification results and dates: [...]
Related incidents, runbooks and reusable lessons: [...]
Close the loop after the document is reviewed
Review publication and action completion are separate milestones. The team can agree that the account is accurate while engineering work remains open. Track both explicitly so a finished document does not quietly become a finished reliability program.
At the next action review, ask the owner to show the acceptance evidence. If a test fails, keep the action open or revise the plan with an explanation. If the service changes before the action ships, confirm that the original risk still exists and that the proposed control still addresses it.
Over time, look for repeated conditions: unrepresentative test workloads, unclear ownership, ineffective rollback checks or missing customer-impact signals. Link related reviews and choose a shared improvement when it addresses several incidents. Do not demand a new process for every individual outage.
Useful measures include overdue high-priority actions, actions with verified acceptance evidence, recurrence of the same failure mechanism and whether the review itself produces usable information. Avoid ranking responders by incident counts or rewarding rapid document closure at the expense of investigation.
Before marking the review complete
- Impact has a defined scope, time window and denominator.
- Important timeline entries point to evidence and use a consistent time basis.
- Facts, explanations and open questions are distinguishable.
- Decision context explains what responders could know at the time.
- Contributing conditions lead to specific, prioritized improvements.
- Each accepted action has an owner, due date, reviewer and acceptance check.
- Useful safeguards and successful response work are recorded too.
- The reviewed document and follow-up work have separate, visible states.
The review earns its place when the next responder has better information, a stronger safeguard or a safer recovery path. Give the document enough structure to make that change possible, then return to the actions until there is evidence that it happened.
For related practical work, see the on-call handoff checklist, the runbook guide and alert grouping without hidden problems. Use the two Google SRE sources above for the broader blameless-review philosophy and the distinction between triggers and contributing conditions.
Want to make this less theoretical?
IncidentRelay gives you schedules, rotations, routing, escalations and alert actions in one open-source, self-hosted package. The pager may still be rude, but at least it will be organized.