On-call culture: the silent tax on your best engineers
We have covered the tools — runbooks to resolve issues, monitoring to catch them early. This week we talk about the people carrying the…
On-call culture: the silent tax on your best engineers
We have covered the tools — runbooks to resolve issues, monitoring to catch them early. This week we talk about the people carrying the pager. Because how you structure on-call says more about your engineering culture than almost anything else. And the cost of getting it wrong rarely shows up where you expect it.
The problem nobody sees coming
On-call does not kill engineers overnight. It does something slower and harder to reverse.
It accumulates.
A missed dinner here. A broken night of sleep there. A weekend half-ruined by an alert that turned out to be nothing. Individually none of it feels like a crisis. Collectively it builds into something your best engineers start calculating quietly, privately against every other option available to them.
By the time someone resigns, the decision was made months earlier. The exit interview will say something polite about career growth. The real reason is that they were exhausted and nobody noticed.
What bad on-call culture actually looks like
It rarely looks obviously broken from the outside. It looks like a team that ships. Engineers who are responsive. A system that mostly stays up.
What it feels like from the inside is different.
The same three people get pulled into every incident because they are the only ones who know enough to help. New engineers are never fully onboarded onto on-call because there is no time to do it properly. Alerts fire constantly! Some critical, many not, and the team has stopped trusting them because the noise-to-signal ratio is too high. Nobody has ever formally reviewed what the on-call experience is like because the answer is uncomfortable.
The engineers who can leave eventually do. The ones who stay are often the ones with fewer options. This is not a technology problem. It is a leadership failure that compounds quietly over time.
The three things that make on-call sustainable
There is no single fix. But every engineering team with a healthy on-call culture has three things in place.
The first is rotation fairness. On-call responsibility is distributed across the team, not concentrated in the people who know the most. Concentration feels efficient in the short term. It is how you burn out your most valuable engineers and create single points of failure in your organization simultaneously.
The second is alert quality. An on-call engineer who gets paged twenty times a week and eight of those are noise has a fundamentally different experience than one who gets paged eight times and all eight require action. Alert fatigue is real, measurable, and almost always ignored until it is too late. Reducing false positives is not a minor housekeeping task. It is a direct investment in the wellbeing of whoever carries the pager.
The third is recovery time. On-call should come with explicit acknowledgment that it is a cost. Time in lieu. Reduced sprint commitments the week after a heavy rotation. Something that signals to engineers that the company sees what they are giving and is not treating it as free. The teams that do this retain people. The teams that do not wonder why their senior engineers keep leaving for companies with “better work-life balance.”
The C-level question that reveals the culture
Here is the question worth asking your engineering leadership this week:
“When did we last ask the team what on-call actually feels like, and what did we do with the answer?”
If nobody has asked, or if the answer produced a list of action items that never moved, you already know where you stand.
What this connects to
Runbooks, monitoring, and on-call culture are not three separate problems. They are the same problem at different layers.
A team without runbooks burns time on known issues. A team without monitoring flies blind until customers complain. A team with broken on-call culture burns through the people who hold everything together.
Fix one without the others and you are patching a system that will keep leaking. The companies that scale without operational chaos are the ones that treat these three things as a connected foundation, and not individual projects to tackle when there is time.
There is a discipline for building and managing that foundation systematically. Next week we will start talking about what it looks like, and why the companies that invest in it early stop fighting the same fires repeatedly.
Subscribe to VisionOps so you don’t miss it.
메타데이터
- post_id
- c09714b60261
- slug
- on-call-culture-the-silent-tax-on-your-best-engineers-c09714b60261
- url
- https://medium.com/@visionops888/on-call-culture-the-silent-tax-on-your-best-engineers-c09714b60261
- canonical_url
- https://medium.com/@visionops888/on-call-culture-the-silent-tax-on-your-best-engineers-c09714b60261
- author_url
- https://medium.com/@visionops888
- status
- ok
- fetched_at
- 2026-08-18 17:22:39