The Context Switching Cost of On-Call Work, and How Teams Reduce It

Image Source: depositphotos.com

Ask an engineer what on-call costs them and most will describe a page at three in the morning. Those nights are real, and they are also comparatively rare on a healthy rota. Teams plan for them, compensate for them, and talk about them openly.

The larger cost is quieter and almost never discussed, because it does not look like an incident. It is what happens to an ordinary Tuesday when you are carrying the pager: work arrives in fragments, nothing deep gets finished, and by Friday you have been busy for five days without being able to say what you built.

That is context switching cost, and it is worth measuring with the same seriousness teams apply to alert volume.

What a switch actually costs

The intuitive model is that an interruption costs as long as the interruption lasts. Glance at a message, spend forty seconds, return. Forty seconds gone.

That model is wrong in a way that compounds. When attention moves from one task to another, part of it stays behind on the first, and the residue takes time to clear. Research on workplace interruptions has put the average time to resume an interrupted task at around 23 minutes. That figure is an average across many contexts rather than a price tag attached to every notification, but the shape of the finding is not in dispute: the resumption is expensive, and the glance is not the cost.

For engineers the reload is often heavier than average, because the thing being reloaded is a mental model of a system. Holding the current state of a distributed trace, a half-understood race condition, or the specific reason this query is slow takes real effort to assemble and almost none to lose.

Six shallow interruptions do not consume forty seconds each. They can consume most of an afternoon.

Why the standard advice does not apply

Almost every piece of focus advice reduces to the same instruction: turn off notifications.

On-call engineers cannot follow it. The interruption channel has to stay open, because staying reachable is the job. Any advice that begins by closing it is describing someone else's week, which is why most focus guidance reads as faintly insulting to people on a rota.

So the useful question is not how to become unreachable. It is which channels genuinely need to reach you, and whether the rest are currently borrowing that privilege.

Separate the channel that must interrupt from the ones that must not

This is where most of the recoverable cost sits.

Ask an engineer coming off a week on call how many times they were paged, and then how many times they were interrupted. The first number is usually small. The second is not. The gap is made up of general chat, direct messages, ticket updates, deployment notifications, standup reminders and email, none of which required a response inside five minutes and all of which arrived on a device kept deliberately open.

The alerting path and the ambient communication path have merged on most teams, usually because both live in the same chat tool. Once they are merged, an engineer cannot close one without closing the other, so they close neither.

Separating them is unglamorous and effective. Route genuine alerts through a path that stays open by design and is reserved for things that need a human now. Treat everything else as interruptible, which means it can be closed during a focus block and read afterwards. The point is not to be less responsive to incidents. It is to stop paying incident-level attention prices for a message about next week's retro.

Alert quality is an attention problem too

Alert tuning is usually argued for on noise reduction grounds, or on the basis of fatigue and missed signals. Both are good arguments and both are well rehearsed.

There is a third one worth adding. Every alert that fires and did not need a human is a context switch spent for nothing, and it is charged at the full rate. The engineer still lost the mental model, still spent the reload, and got no incident in exchange.

Framed that way, an alert that fires forty times a month and requires action twice is not merely noisy. It is consuming a meaningful slice of the team's capacity for deep work, and the cost never appears in any dashboard because nothing broke.

Protect the weeks when nobody is paging you

Most teams that take on-call seriously put their effort into making the on-call week survivable: sensible handovers, follow-the-sun where headcount allows, compensating time, clear escalation.

Far fewer look at the other weeks. The assumption is that the rota rotates and the engineer recovers. That only holds if the weeks off the rota are structurally different, and often they are not. If the non-on-call week is fragmented by the same ambient interruptions, there is no recovery period, only a period with fewer pages in it.

The practical move is to treat protected time off the rota as part of the on-call design rather than a personal productivity concern. A few blocks a week where an engineer is not reachable by anything except a genuine incident is not a perk; it is what makes the rota sustainable across a year.

Individual tooling helps here, and the distinction that matters is who configures it. Tools like DigitalZen work on an allow-only basis, where the engineer names the two or three things a task needs and closes the rest for a defined window, with the alerting path deliberately left open. This is adopted by the engineer rather than imposed on them, which is the line between a focus tool and a monitoring tool. Nothing here should involve visibility into how anyone spends their time; that is a different product category and a worse idea.

It also tends to work better when the block is decided in advance, since pre-committing a schedule while you are still calm produces more sensible rules than deciding mid-afternoon with a deadline approaching.

Three changes worth making first

If this is worth acting on, the order matters more than the completeness.

Measure interruptions, not just pages. Ask the outgoing engineer at handover how many times they were pulled away and by what. Two weeks of that will tell you where the cost is, and it usually is not the pager.

Split the alerting channel from the chat channel. If an engineer cannot close general communications without risking a missed incident, they will close nothing.

Give the non-on-call weeks a shape. Protected blocks, agreed by the team, honoured by the team. This costs nothing and is the single change most likely to survive contact with a busy quarter.

None of this reduces how quickly anyone responds to a real incident. It reduces how much the team pays for everything that was not one.