Many runbooks begin as excellent intentions and end as a wiki page last updated by someone whose account is now disabled. The title looks promising, the architecture diagram is from three platforms ago, and the first instruction is "check the logs" — a sentence with roughly the operational value of "investigate the mystery".
A strong on-call runbook is a decision tool. It should move the responder from alert to verified impact, then to a safe mitigation or an explicit escalation. The target is not complete knowledge. The target is useful action under pressure.
The 30-second runbook test
Open the runbook as if you have never seen the service. Within 30 seconds, can you answer:
- Which service and environment does this apply to?
- What user or business impact does the alert imply?
- What are the first two safe checks?
- What action may reduce impact?
- What must not be changed without approval?
- When and to whom should the incident escalate?
If the responder must read four pages of history before finding the first command, the runbook is documentation with an adventurous filename.
Start from an alert contract
Write one runbook for a meaningful operational condition, not for a vague component. "Kubernetes runbook" is an encyclopedia project. "Checkout API high error rate in production" has a trigger, an impact and a set of decisions.
| Contract field | Question | Example |
|---|---|---|
| Trigger | Exactly which alert or condition opens this runbook? | HighErrorRate for checkout-api in production |
| Urgency | How quickly must a human act? | Acknowledge within 10 minutes |
| Expected impact | What may users observe? | Checkout requests fail or time out |
| Owner | Which team can perform the actions? | Payments SRE |
| Exit | What proves the condition is controlled? | Error rate below 1% for 10 minutes |
This contract also exposes bad alerts. If nobody can define the required action or exit condition, the signal may belong in a dashboard or ticket instead of waking a human.
Put the critical path at the top
The first screen should contain the information needed in the first five minutes:
- Scope: service, environment, region and alert name.
- Impact: what customers or dependencies may experience.
- Safety warning: actions that risk data loss, failover or wider impact.
- Quick links: dashboard, logs, recent deploys, traces and service status.
- First checks: two or three ordered, read-only validations.
- Decision: a small if/then tree leading to mitigation or escalation.
Background architecture, historical incidents and detailed theory can live below the critical path or in linked documentation. The responder should not scroll past "A brief history of distributed consensus" while checkout is unavailable.
Use ordered checks with expected results
"Check metrics" and "look at logs" are not steps. Name the view, time window, query and meaning of the result.
| Weak instruction | Operational instruction |
|---|---|
| Check the dashboard. | Open Checkout Overview, set window to last 30 minutes, compare request rate and 5xx rate by region. |
| Check recent changes. | Open Deployments, filter checkout-api, note releases completed within 30 minutes before the alert. |
| Check logs. | Open the saved "checkout 5xx" query and group errors by exception name and upstream dependency. |
| Verify health. | Compare the external synthetic check with internal health endpoints; disagreement suggests ingress or dependency failure. |
Every check should say what to do with the result. A beautiful graph without a decision is incident-themed wallpaper.
Prefer read-only commands first
The first commands should collect state without changing it. For example:
# Confirm deployment and rollout state
kubectl -n production get deploy checkout-api
kubectl -n production rollout status deploy/checkout-api
# Inspect recent application events
kubectl -n production get events --sort-by=.lastTimestamp
# Verify the public health path
curl -fsS https://checkout.example.com/health
These are examples, not universal commands. A real runbook must name the correct cluster context, namespace, account and expected output. "Run kubectl" without a context check is a short story about how staging survived while production got weirder.
Put guardrails around every write action
A mitigation that changes production needs more than a command block. Include:
- When to use it: the evidence that justifies the action.
- Expected effect: which metric should improve and how quickly.
- Risk: sessions dropped, backlog growth, reduced redundancy or data concerns.
- Approval: whether the primary can act alone or needs a second responder.
- Verification: exact checks after the change.
- Rollback: how to reverse it and when reversal becomes unsafe.
Do not make a sleep-deprived responder invent the rollback for a command the runbook told them to copy.
Prefer a tested automation job, deployment rollback control or script with validation over a long sequence of manual commands. If a command can delete data, change multiple regions or operate without a scope limit, do not present it as a casual copy-paste step.
Write decisions as branches
Long prose hides choices. A small decision tree makes ownership explicit:
Is customer-facing error rate above the incident threshold?
no -> verify alert recovery, watch for 10 minutes, then resolve
yes -> was a checkout-api deploy completed in the last 30 minutes?
yes -> follow approved rollback procedure
no -> is one upstream dependency failing?
yes -> open dependency runbook and enable approved fallback
no -> page secondary and start broader incident triage
Keep branches based on evidence available during the incident. "If the root cause is network entropy" may be scientifically exciting but is not a branch a responder can evaluate in five minutes.
Define escalation triggers before the incident
"Escalate if needed" transfers the hardest decision to the most tired person. Use observable triggers:
- no safe mitigation identified within 10 minutes;
- data integrity or security impact is suspected;
- two regions or multiple critical services are affected;
- rollback fails or increases impact;
- the responder lacks permissions for the required action;
- customer impact crosses the P1 threshold;
- the incident continues beyond the team's response objective.
Name the destination and action: page the secondary rotation, activate the escalation policy, request a database specialist, or assign an incident commander. A contact named "ask around in Slack" is not a dependency with an SLA.
Step 7: "Contact Dave."
Dave left in 2023. The incident has now become a séance.
A compact example: Checkout API high error rate
RUNBOOK: Checkout API high error rate
Owner: Payments SRE
Applies to: production / checkout-api / HighErrorRate
Last tested: 2026-08-01
IMPACT
Customers may be unable to complete checkout.
DO NOT
- restart the database;
- disable payment validation;
- fail over regions without the database lead.
FIRST 5 MINUTES
1. ACK the alert and confirm incident priority.
2. Open Checkout Overview and compare 5xx by region.
3. Check deploys completed in the previous 30 minutes.
4. Open saved error-log query and group by exception.
DECISION
- Recent deploy + matching error spike:
run the approved rollback workflow, then verify for 10 minutes.
- One payment provider failing:
follow Provider Failure runbook and enable approved fallback.
- Multiple services or unknown cause:
page secondary and start P1 coordination.
RECOVERY
- 5xx below 1% for 10 minutes;
- synthetic checkout succeeds in every active region;
- queue backlog is stable or falling;
- no unresolved data-integrity warning.
FOLLOW-UP
Link the incident, timeline, change and owner for permanent repair.
The thresholds and actions are illustrative. Replace them with values validated for the service. The useful pattern is stable: scope, impact, prohibition, ordered checks, decision, recovery and follow-up.
Connect the runbook to the alert
A perfect runbook is useless if the responder cannot find it. Attach operational context where the incident begins.
In IncidentRelay, a service can contain stable links to dashboards, logs and traces plus one or more runbooks. A runbook with empty matchers can apply to every alert for the service; alert-specific matchers can select instructions for conditions such as HighErrorRate in production.
{
"alertname": "HighErrorRate",
"environment": "production",
"service": "checkout-api"
}
When the incoming alert is attached to the service, matching runbooks and service links can appear in alert details and supported notifications. This keeps the stable service context separate from the monitoring rule that happened to detect the problem.
See Services for the service and runbook model. Teams can also define service standards that check whether important services have owners, links and runbooks.
Avoid matcher traps
Alert-specific runbooks are only useful when their matchers reflect stored alert data. Use stable labels and verify spelling, case and environment values. Conditions use AND semantics, so every field must match.
| Problem | Symptom | Fix |
|---|---|---|
Matcher uses env=prod, alert stores environment=production. | The generic runbook appears, but the specific one does not. | Inspect stored labels and standardize the integration. |
| Runbook matches only an instance name. | Autoscaling replaces the host and instructions disappear. | Match stable service and alert labels. |
| Every alert gets ten runbooks. | The responder must guess which one matters. | Narrow specific matchers and keep one concise generic fallback. |
| Runbook URL requires a separate broken identity system. | The link exists but cannot be opened during the outage. | Test incident-time access and provide safe fallback context. |
Test with someone who did not write it
The author knows what every ambiguous sentence was supposed to mean. A new responder does not. Use three levels of testing:
- Tabletop: give the alert to another engineer and ask them to narrate each decision without changing production.
- Staging exercise: create the condition in a safe environment and execute the checks and mitigation.
- Game day: send a realistic alert through routing and require the on-call responder to use the runbook, ACK and escalation path.
Observe time to first useful check, broken links, permission failures, commands that produce unexpected output, unclear branches and missing stop conditions. Our on-call game day guide provides a ready-to-run exercise format.
Review from events, not calendar guilt
A quarterly review date is useful, but operational events are better triggers. Review the runbook when:
- the alert fires and the responder deviates from the documented steps;
- service architecture, deployment tooling or ownership changes;
- a dashboard, query, credential flow or escalation policy changes;
- a mitigation fails or creates unexpected impact;
- a game day finds ambiguity;
- a new responder cannot complete the tabletop exercise.
Store an owner, last reviewed date and last tested date. "Reviewed" means links and wording were inspected. "Tested" means a human followed the steps against a realistic condition. These are different levels of confidence.
Measure whether it helps
Useful runbook metrics are modest and outcome-oriented:
- percentage of paging alerts with a reachable relevant runbook;
- time from ACK to first diagnostic action;
- incidents where responders used an undocumented workaround;
- broken-link or permission failures found during tests;
- successful mitigations versus escalations caused by unclear guidance;
- runbooks tested after their last material service change.
Do not reward word count or number of runbooks. One maintained decision path beats twenty pages of archaeology.
The copy-paste template
# [Service] — [Alert/condition]
Owner:
Applies to:
Last reviewed:
Last tested:
Required access:
## Impact
[What users and dependencies may experience]
## Safety / do not
[Dangerous actions, approvals and data risks]
## Quick links
- Dashboard:
- Logs:
- Traces:
- Recent changes:
## First five minutes
1. [Read-only check + expected result]
2. [Read-only check + expected result]
3. [Read-only check + expected result]
## Decision tree
- If [evidence], then [safe action].
- If [evidence], then [other runbook/action].
- If unclear after [time], escalate to [target].
## Mitigation
[When, risk, approval, action, verification and rollback]
## Escalate when
[Observable triggers and exact destination]
## Recovery criteria
[Metrics, duration and user validation]
## Follow-up evidence
[Incident link, timeline, change, owner and repair task]
Definition of done
The runbook is ready when a qualified engineer who did not write it can open it from the alert, understand impact, complete the first safe checks, choose a branch, identify prohibited actions, escalate correctly and verify recovery. All links work with incident-time permissions, and at least one realistic exercise has followed the steps.
Keep it short enough to scan, specific enough to act on and honest enough to say "stop and escalate" when the safe answer is not known. At 3 AM, clarity is a reliability feature. So is not having to contact Dave.
Want to make this less theoretical?
IncidentRelay gives you schedules, rotations, routing, escalations and alert actions in one open-source, self-hosted package. The pager may still be rude, but at least it will be organized.