Skip to the content
Work

The Maths That Lets a Six-Person Team Run On-Call Without Burning Out

A workable on-call rotation is not a schedule on a wall. It is a promise that someone will answer a page within minutes, even at 2am, and that the waking will be rare enough to keep judgment intact. The Google SRE Workbook, last updated in 2026, defines the hard limit: two distinct paging incidents per 12-hour shift.

6 min read

An outdoor signal bell on its bracket, rung to call somebody out
An outdoor signal bell on its bracket, rung to call somebody out. Photo: Wolfmann · Wikimedia Commons · CC BY-SA 4.0

Cross that line and the on-call engineer's error rate rises with their fatigue. For a six-person team, that limit becomes an arithmetic problem. Solve it honestly, or the rotation collapses into resentment and missed pages.

What the Rotation Actually Buys

PagerDuty's on-call guide, states the expectation plainly: the person on-call must be able to respond at 2am, and escalation to the wider team happens within five minutes. This is not a courtesy. It is a service-level agreement with whoever depends on the system overnight. The Google SRE Workbook adds the sharper constraint: every alert that wakes a human must be immediately actionable. There must be something the recipient is expected to do right then, not document for Monday. PagerDuty's alerting principles echo this: "anything that wakes a human in the middle of the night should be immediately human actionable." An alert that merely informs, or that requires investigation before action, has no business paging at night.

This distinction matters because many teams treat their on-call roster as a notification dump. PagerDuty's documentation notes that if an event arrives with no on-call user assigned, the event is dropped entirely. The system assumes intentionality. The team must match that intentionality in what it labels as paging-worthy.

The Arithmetic of Six People and Seven Days

Google SRE defines pager load as the number of paging incidents per shift length, typically measured daily or weekly. The ceiling is two incidents per 12 hours. A week-long rotation, as SquareOps noted in a 2026 framework, can mean seven consecutive nights of interrupted sleep. With six people, the rotation cycles every six weeks. If the system pages twice per night, the math is brutal: 14 incidents per week, 84 per rotation, 728 per person per year.

SquareOps warns that more than two pages per 12-hour shift is unsustainable. The framework also notes that fewer than 30% actionable alerts signals alert fatigue. A six-person team can survive only if the actual page volume sits well below the theoretical maximum. Incident.io's 2026 best-practices guide suggests 2–3 actionable incidents per shift as a workable range, but the upper bound assumes optimal conditions: trained responders, clear runbooks, no flapping alerts. Most teams are not optimal.

The rotation length compounds the problem. SquareOps recommends a minimum five-person rotation to achieve on-call every five weeks. A two-person rotation, by contrast, means on-call every other week with zero flexibility for illness or vacation. Six people is barely above that threshold. One departure drops the team to a five-week cycle with no slack.

Two Numbers That Matter More Than Page Count

Raw volume misleads. The Google SRE Workbook emphasizes signal-to-noise ratio: low signal raises alert fatigue risk. SquareOps quantifies this: fewer than 30% actionable pages suggests a systemic problem. The actionable rate matters because non-actionable pages still fracture sleep and erode trust in the system.

A team of six should track two interrupt-load measures. First: pages outside working hours per person per week. This captures the sleep cost directly. Second: the share of pages that were actionable, meaning the responder could and did take immediate corrective action. A page that merely acknowledges a threshold breach, or that queues work for the next business day, drags both numbers in the wrong direction.

Incident.io's 2–3 actionable incidents benchmark assumes the responder can act. If your runbooks are incomplete or your permissions are misconfigured, the actionable rate drops even when the alert is technically correct. The metric exposes that gap.

Separating Wake-Ups from Waiting

Not everything that needs human attention needs it now. PagerDuty's support-hours configuration distinguishes high-urgency paging during support hours from low-urgency incidents outside those hours. The low-urgency rules can be set quiet: no phone call, no 2am vibration, perhaps a logged ticket for morning review. This is where many teams fail. They conflate "needs a human" with "needs a human immediately."

PagerDuty's documentation shows how to pair support-hours behavior with responder notification rules. Outside-hours incidents drop to low urgency; low-urgency rules stay silent. The alert still exists. It simply waits for working hours. This requires discipline in how services are tagged and how thresholds are set. An alert that fires at 3am for a condition that has no customer impact until 9am is misclassified as high-urgency. Google SRE's immediately actionable rule is the filter: if the action can wait, the alert should not page.

There is also the trap of the dropped alert. PagerDuty notes that events with no on-call user are dropped entirely. This is not a safety feature to rely upon. It is a failure mode that creates silent gaps in coverage. The roster must be maintained with the same rigor as the code.

When the Maths Fails: Escalation, Handover, and Escape Hatches

PagerDuty's five-minute escalation window is tight. It assumes the on-call person is awake, connected, and able to engage. SquareOps warns that a seven-day rotation means an entire week of interrupted sleep. The framework recommends reducing alert noise and shortening rotations if needed, though it offers no formal threshold for when shortening becomes mandatory.

The practical handover matters. PagerDuty's 2am expectation means the outgoing on-call engineer may be handing off incidents that started minutes before the rotation change. The incoming engineer inherits partial context and fractured sleep. Secondary sources describe handoff practices, but no authoritative template exists for what belongs in a handover artifact. Teams improvise: open incident links, pending follow-ups, alerts that fired but auto-resolved. The improvisation itself creates risk.

When the page load exceeds what six people can sustain, two structural fixes appear in practice literature, though neither is formally mandated by a standards body. One is a follow-the-sun partner: another team in a complementary timezone shares the rotation, cutting each team's overnight burden. The other is narrowing the paging scope: removing services from the on-call roster, accepting longer recovery times for less critical systems, and converting wake-up alerts to Monday-morning reports. Both require organizational buy-in. Neither is free.

The Rotation Must Be Revisited After Real Data

Google SRE's framing of pager load and signal-to-noise assumes continuous measurement. The incident.io benchmark of 2–3 actionable incidents per shift is a starting point, not a permanent guarantee. SquareOps recommends reviewing continuously, not on a fixed calendar. The relevant data accumulates only after several rotation cycles: nights disturbed per person, actionable rate, escalation frequency, incidents that auto-resolved before human intervention.

A six-person team can survive on-call if the arithmetic is honest. Two pages per 12-hour shift is the ceiling, not the target. The actionable rate must stay above 70% to prevent alert fatigue. Support-hours configuration must keep non-urgent work out of night hours. The rotation must shorten or expand based on actual paging data, not hope. The system survives only if the team keeps asking whether each 2am wake-up was truly necessary, and has the authority to delete the alert when the answer is no.

Sources

  1. Google SRE Workbook, “What it Means Being On-Call?” — sre.google, 2026-10-02
  2. PagerDuty, “Alerting Principles” — response.pagerduty.com, 2026-10-02
  3. PagerDuty Knowledge Base, “Configurable Service Settings” — support.pagerduty.com, 2026-10-08
  4. PagerDuty Knowledge Base — support.pagerduty.com, 2026-10-02
  5. PagerDuty, “The On Call” — pagerduty.com, 2025-04-15
  6. incident.io blog, “On-call best practices: handoffs, schedules, and alert fatigue” — incident.io, 2026-02-27
  7. SquareOps, “Reducing Alert Fatigue: A Practical On-Call Framework” — squareops.com, 2026-07-16

More from Work & Leadership

Section index
Work

How to apply, as the firm described it

Join a team of professionals dedicated to one another’s success. Contact us to learn about our job opportunities and to find out why the firm has been consistently identified as one of the…

29 Sep 2026
2 min

Work

Careers at a Pittsburgh promotions firm

Innovation is the key to success in business. We believe this so strongly at the firm that we consider team development to be one of our highest priorities. By building a thoroughly trained…

25 Sep 2026
2 min

Work

Replacing resolutions with habits

It seems that every year, the masses make new resolutions and vow to keep them, only to falter a few weeks into their commitments. At the firm, we came across one study indicating that…

31 Aug 2026
2 min

Independent trade desk. We take no commission on anything we describe and run no affiliate programme of our own. Every figure on this page names the standard, filing or organisation it comes from; where a number could not be verified the page says so. How we work and how we correct. Reviewed:

Cookies, and what this site stores. The Dispatch sets no advertising or analytics cookies and loads no third-party tracker. Closing this notice writes one key — icd-notice — into your browser’s local storage, so that the notice does not return. Nothing else is kept. The one thing a page here sends onward is what a reader types into the form on the contact page, and that is described before the form is used. What the policy says.