← Blog On-call capacity

How many engineers do you actually need for 24/7 on-call?

Two engineers can fill every square on a 24/7 calendar. That does not mean two engineers can sustainably operate it. Let us calculate the difference between technically covered, operationally resilient and quietly setting the team on fire.

2026-07-28 11 min read 24/7 on-call SRE staffing on-call rotation

There is no magic headcount that makes 24/7 on-call safe. A database team receiving one actionable page per month has a different staffing problem from a platform team receiving twelve every night. But there is useful arithmetic, and it is much better than deciding that "everyone will share the load" while everyone slowly updates their résumé.

A sensible staffing estimate combines three questions:

  1. Calendar capacity: can every primary and backup shift be assigned?
  2. Response capacity: can the assigned people diagnose and act within the required time?
  3. Recovery capacity: can the team absorb vacations, illness, incidents and rest after a bad night?
IncidentRelay on-call calendar
A full calendar proves that names fit into boxes. We still need to ask whether the humans inside those boxes remain functional.

Count qualified responders, not team members

Start with the number of people who can independently take the first useful action for the service. A qualified responder has:

  • working access to production, dashboards, logs and communication channels;
  • enough system knowledge to assess impact and choose a safe first action;
  • permission and authority to mitigate, roll back or escalate;
  • a tested device and notification path;
  • an understood handoff, acknowledgement and escalation process.

If the team has eight people but only three can restart the critical service or access the production database, the rotation has three qualified responders. The other five may be training capacity, which is valuable, but calendars do not convert optimism into permissions.

Calculate the primary rotation baseline

For weekly primary shifts, the simplest annual estimate is:

primary weeks per responder = 52 / qualified responders
Qualified respondersPrimary weeks/person/yearAverage spacingPlanning signal
226Every other weekExtremely fragile
317.3Every third weekHeavy and absence-sensitive
413Every fourth weekPossible at low page volume
510.4Every fifth weekWorkable baseline for many teams
68.7Every sixth weekMore room for recovery and training
86.5Every eighth weekComfortable only if skills are genuinely shared

This table is not a law. Four people with almost no nighttime pages may operate comfortably; eight people receiving hundreds of unactionable pages will still hate the pager. It is a starting point for asking better questions.

For a small team, four qualified responders is often a fragile lower bound; five or six gives a more credible starting point. Page load and backup requirements can move that number sharply upward.

Add the availability factor

People take vacations, attend training, become ill, leave the company and occasionally need a weekend that is not a character-building exercise. Model the fraction of the year each responder is realistically available for rotation.

effective responders = qualified responders × availability factor

planning primary load = 52 / effective responders

An availability factor of 0.80 does not mean people work only 80% of the year. It reserves capacity for planned leave, training, recovery and ordinary organizational friction. Choose the factor from your actual calendar and policies rather than borrowing 0.80 because it looks satisfyingly round.

TeamQualifiedAvailability factorEffective capacityPlanning load
Team A40.803.216.3 primary weeks/person-equivalent
Team B50.804.013 primary weeks/person-equivalent
Team C60.824.910.6 primary weeks/person-equivalent

The "person-equivalent" wording matters: individuals are not decimal engineers. The number estimates the burden that remaining available responders must absorb when someone is away.

Secondary coverage is another duty line

If critical alerts require a separate backup, you are scheduling 104 duty-weeks per year: 52 primary and 52 secondary. The secondary may receive fewer pages, but they still have constraints: they must remain reachable, keep access nearby and avoid activities incompatible with taking over.

annual duty-weeks = 52 × duty lines

primary + secondary = 52 × 2 = 104 duty-weeks

With six people sharing both roles evenly, each person averages about 8.7 primary weeks and 8.7 secondary weeks per year before absence cover. That is 17.4 weeks attached to some form of duty. Do not describe the secondary as "free" merely because their phone screams less often.

Offset the primary and secondary rotations so they cannot resolve to the same person. Also define what secondary means:

  • must they acknowledge within five minutes?
  • must they have a laptop and production access immediately?
  • do they take ownership or only help the primary?
  • when can they be contacted for non-critical questions?
Capacity planning, enterprise edition

"The manager is always the backup."
Have we informed the manager, the calendar, or reality?

Measure interruptions, not only assigned weeks

Scheduled time tells you how often someone carries the pager. Alert history tells you what that duty costs. Track at least:

  • actionable pages per primary shift;
  • after-hours pages and distinct nighttime wake-ups;
  • time spent responding, not only time to ACK;
  • escalations to the secondary;
  • concurrent incidents and multi-hour incidents;
  • recovery time or next-day work displaced by an incident.

One incident that sends fifteen grouped notifications is different from fifteen independent incidents, and three pages at 02:00, 03:30 and 05:00 are worse than three pages during lunch. Use event timelines, not just a total counter.

A simple per-person estimate is:

expected after-hours pages/person/month
  = team after-hours pages/month / qualified responders

Then inspect distribution and bursts. An average of two can mean exactly two each, or one engineer receiving ten during the week the database developed opinions.

A worked example: five engineers

Suppose five qualified engineers operate one production service with weekly primary and secondary rotations:

Qualified responders5
Availability factor0.80
Effective planning capacity4.0
Primary planning load13 weeks/person-equivalent/year
Secondary planning loadAnother 13 weeks/person-equivalent/year
After-hours pages18 per month across the team
Average after-hours pages3.6 per responder/month before burst adjustment

The calendar can be filled, but the system is warm: roughly half the year may involve primary or secondary constraints after allowing for absences, and the team receives meaningful nighttime interruption. The next step is not automatically "hire exactly one person". First inspect whether those 18 pages are actionable, grouped and owned. Then compare the remaining load with the team's recovery policy and response objective.

If most pages come from two noisy rules, fix the rules. If they represent real incidents and business risk requires faster parallel response, add qualified capacity. Staffing should not subsidize monitoring defects forever, and alert tuning should not hide a genuine capacity shortage.

Run the one-person-away test

Remove any one responder from the model for two weeks. Then ask:

  • Can primary and secondary still be different qualified people?
  • Can planned overrides cover the absence without repeated double shifts?
  • Does someone lose all meaningful recovery time between duties?
  • Is a specialist required for one service, making the general rotation fictional?
  • Can the remaining team handle a long incident and normal project work?

Repeat it for each person, especially the most experienced responder. If removing one name collapses the schedule or removes access to the only safe mitigation, you have a bus-factor problem disguised as a rotation.

Follow-the-sun changes the arithmetic

Follow-the-sun can reduce night work by assigning regional layers during local waking hours. It does not reduce the need for qualified coverage; it distributes it.

Two regions covering twelve hours each still create long duty windows and difficult daylight-saving boundaries. Three regions can create cleaner eight-hour windows, but each region needs enough trained people to survive leave and turnover. Three regions with one expert each is not resilience. It is a globally distributed set of single points of failure.

Plan overlap for handoff, use explicit layer timezones, and decide who owns an incident that crosses the boundary. IncidentRelay rotation layers can model business hours, nights, weekends and regional windows; higher-priority active layers win when windows overlap. Verify the final result in the calendar across daylight-saving transitions.

Know the signs that staffing is too thin

  • Vacation approval depends on finding a personal swap.
  • The same expert joins incidents even when someone else is primary.
  • Primary and secondary frequently collapse onto the same person.
  • Responders work normal project days after repeated nighttime incidents.
  • New engineers are added to the rotation before they are qualified.
  • Alert thresholds are weakened to protect an overloaded team rather than because risk changed.
  • The calendar is filled only through long-lived manual overrides.
  • Nobody can explain how coverage survives one resignation.

These are system signals, not evidence that people need to become more heroic. Heroism is useful during a rare incident and a terrible recurring staffing model.

Use more than one lever

LeverWhat it improvesTrade-off
Hire or transfer qualified respondersReduces duty frequency and specialist risk.Slow and requires onboarding.
Cross-train existing engineersTurns nominal headcount into qualified capacity.Consumes senior time before it returns capacity.
Reduce paging noiseLowers interruption cost immediately.Requires engineering ownership, not mass silencing.
Automate safe remediationShortens or removes repetitive response work.Automation needs guardrails and testing.
Limit coverage hours or severityConcentrates paging on risks that justify it.The business accepts slower response elsewhere.
Add regional coverageReduces nighttime duty.Needs regional skills and strong handoffs.

Often the right answer combines them: remove five non-actionable pages, train two responders, and hire one engineer for a genuine coverage gap. Capacity planning is allowed to have more than one cell in the spreadsheet.

Model the result in IncidentRelay

  1. Create one rotation for each distinct duty line, such as primary and secondary.
  2. Use ordered members and explicit start dates for qualified participants.
  3. Use layers and restrictions for recurring regional or business-hours coverage.
  4. Use short, reasoned overrides for leave and swaps.
  5. Connect routes to the correct rotation or escalation policy.
  6. Open On-call Health and inspect future gaps, inactive members and single-member warnings.
  7. Review the next 90 days in the calendar, including holidays and timezone transitions.
  8. Run an on-call game day to verify delivery, ACK and backup escalation.

The calendar should express the staffing decision, not conceal it. If health diagnostics show a single active member, the fix is not to admire the yellow icon until morale improves.

The capacity worksheet

Bring these inputs to the staffing conversation:

Qualified responders:             ____
Availability factor:              ____
Required duty lines:              primary / secondary / other
Primary weeks/person/year:        52 / effective responders
Total duty-weeks/person/year:     (52 × duty lines) / effective responders
After-hours pages/month:          ____
Distinct night wake-ups/month:    ____
Secondary escalations/month:      ____
Median and p90 incident duration: ____
Longest planned absence:          ____
Single-person specialist risks:   ____
Required acknowledgement time:    ____

Finally, simulate one person away, one bad incident week and one simultaneous vacation. If the model works only when everyone is healthy, available and receiving evenly spaced alerts, it is a demo environment.

So, what is the number?

For one low-volume 24/7 primary rotation, five or six qualified responders is a practical planning baseline for many teams. Four can work but is sensitive to absences. Two or three is usually a temporary risk position, not a sustainable destination. A real secondary duty line, frequent night pages, specialist dependencies or regional coverage can require substantially more.

Do not take that paragraph to finance without the worksheet. The defensible number comes from coverage lines, actual availability, incident load, response objectives and recovery needs. The goal is not merely to put a name in every calendar square. It is to build a rotation that still works after a vacation, a bad night and the departure of the person who knew where all the dashboards were buried.

Want to make this less theoretical?

IncidentRelay gives you schedules, rotations, routing, escalations and alert actions in one open-source, self-hosted package. The pager may still be rude, but at least it will be organized.