← Back to list

We Rewrote 8% of the System in Rust. Pager Volume Dropped More Than Any Infrastructure Upgrade.

For nearly a year, our reliability roadmap looked exactly like you’d expect. Bigger cache clusters. More aggressive autoscaling. Additional…

Computer Architect · 2026-06-04 13:01 · 0 claps · 4.1 min read paywalled
#programming #technology #software-development #rust #design-systems
Open on Medium ↗
Wiki topics: PRD · Product Design 💻 · Programming

We Rewrote 8% of the System in Rust. Pager Volume Dropped More Than Any Infrastructure Upgrade.

For nearly a year, our reliability roadmap looked exactly like you’d expect. Bigger cache clusters. More aggressive autoscaling. Additional observability tooling. Query optimization projects.

Load-balancer tuning. Database upgrades. Every quarter produced measurable improvements somewhere. CPU went down. Latency improved. Infrastructure costs stabilized.

Yet one metric barely moved. Pager volume. The system was getting faster. Engineers were not getting more sleep. That discrepancy became difficult to ignore.

The Number Nobody Wanted to Present

During one reliability review, we plotted six months of operational metrics on a single dashboard. Most graphs looked healthy.

+-----------------------------------+------------------+---------+ 
| Metric                       | 6 Months Earlier | Current | 
+-----------------------------------+------------------+---------+ 
| p95 Latency                     | 214ms       | 147ms | 
| Database CPU                    | 71%         | 48% | 
| Cache Hit Rate                  | 83%         | 94% | 
| Infrastructure Cost per Request | Down 18%    | | 
| Pager Alerts per Week           | 137         | 124 | 
+-----------------------------------+------------------+---------+

The room became unusually quiet. We had spent hundreds of engineering hours improving nearly every visible metric. Pager volume had barely changed. The graph exposed something uncomfortable.

Users experienced a faster system. Operators experienced almost the same amount of pain. Those are not always the same problem.

Every Incident Looked Different

The frustrating part was that postmortems rarely pointed to the same root cause.

One week:

  • request timeout cascade

Another:

  • worker backlog explosion

Then:

  • retry amplification

Then:

  • queue saturation

Then:

  • intermittent memory exhaustion

Then:

  • downstream overload

Every incident appeared unique. Every investigation ended somewhere different. No single service looked catastrophic. No single database query dominated incident reports. No obvious bottleneck appeared.

The architecture looked healthy enough. Yet something kept converting minor disturbances into pages.

The Pattern We Missed

Eventually we stopped grouping incidents by symptom. We grouped them by request path. That changed everything. A surprising percentage of pages touched the same portion of the architecture. Not the largest service.

Not the busiest service. Not the most expensive service. A relatively small processing layer sitting directly on the critical path. Only about 8% of total application code lived there.

But nearly every request passed through it. When load increased, small inefficiencies accumulated. When downstream systems slowed, retries accumulated. When memory pressure increased, latency accumulated.

The failures were different. The amplification point was the same.

Infrastructure Was Treating the Symptoms

Most of our reliability work targeted consequences. Cache improvements reduced pressure. Autoscaling added capacity. Database tuning improved throughput. Observability accelerated diagnosis. All valuable work.

None addressed the amplification mechanism itself. The critical path still transformed small disturbances into larger operational events. Infrastructure upgrades were effectively making the blast radius smaller. They were not eliminating the source of amplification.

The Rewrite Nobody Wanted

The team was reluctant to rewrite anything. For good reason. Most rewrites fail to justify themselves. The proposal wasn’t:

“Let’s migrate everything.”

The proposal was:

“Let’s replace only the highest-leverage execution path.”

Roughly 8% of the codebase. No feature changes. No architectural redesign. No platform migration. Just the request-processing layer that appeared repeatedly across incident timelines. The goal wasn’t speed. The goal was predictability.

What We Found Under Load

Synthetic benchmarks looked good before the rewrite. Production behavior looked different. Under sustained concurrency, the service exhibited behaviors that rarely appeared in isolated testing. Small allocations accumulated.

Retry paths generated bursts of work. Temporary memory growth increased tail latency. Tail latency triggered more retries. Retries created more pressure. Pressure created more tail latency.

The loop fed itself. The system didn’t fail immediately. It gradually destabilized. That’s often worse. Immediate failures are obvious. Slow amplification can survive for months.

The Smallest Deployment of the Year

The deployment itself was uneventful. No launch announcement. No migration celebration. No all-hands coordination. Traffic shifted gradually. The first thing engineers noticed wasn’t latency.

It was silence. Several recurring alert patterns simply stopped appearing. The incident dashboard looked unfamiliar. Not because performance exploded upward. Because operational noise disappeared.

The Graph We Kept Returning To

After six weeks we reviewed the results. Latency improved. CPU improved slightly. Memory consumption became more predictable. But the most surprising graph wasn’t technical. It was operational.

+-------------------------+---------------+-----------------------+ 
| Metric                  | Before        | After | 
+-------------------------+---------------+-----------------------+ 
| Weekly Pages            | 124           | 61 | 
| Timeout Alerts          | 100% Baseline | 41% | 
| Retry Storm Incidents   | Frequent      | Rare | 
| Queue Saturation Events | Recurring     | Occasional | 
| Critical Escalations    | Regular       | Significantly Reduced | 
+-------------------------+---------------+-----------------------+

None of these numbers individually justified a conference talk. Together they explained why engineers suddenly trusted the system more.

Reliability Is an Amplification Problem

The lesson wasn’t “Rust is magic.” That would be the wrong conclusion. The lesson was that reliability bottlenecks often hide in amplification layers. A component doesn’t need to consume the most CPU.

It doesn’t need to allocate the most memory. It doesn’t need to generate the most traffic. It only needs to sit in a position where small problems become larger ones. Those components create disproportionate operational impact.

Removing them changes incident frequency far more than improving average-case performance.

What We Gave Up

The rewrite wasn’t free. Build times increased. Developer onboarding became slightly harder. The implementation demanded more discipline. Some engineers preferred the ergonomics of the previous stack.

Those concerns were legitimate. The project would have been difficult to justify if the outcome had been a 5% latency improvement. A modest benchmark gain would not have paid back the complexity. Operational stability did.

That’s a different equation.

Looking Back

When people ask about the project, they usually ask for latency numbers. Or throughput numbers. Or memory numbers. Those metrics improved. But they’re not what most engineers remember.

What people remember is opening the incident dashboard and realizing familiar alerts were gone. Infrastructure upgrades helped us manage growth. The rewrite changed how failures propagated. Only about 8% of the system changed.

Yet that small section happened to sit where reliability was won or lost. In retrospect, the most valuable optimization wasn’t making requests faster. It was making minor problems stay minor.


메타데이터
post_id
fce168e35536
slug
we-rewrote-8-of-the-system-in-rust-pager-volume-dropped-more-than-any-infrastructure-upgrade-fce168e35536
url
https://medium.com/@pmLearners/we-rewrote-8-of-the-system-in-rust-pager-volume-dropped-more-than-any-infrastructure-upgrade-fce168e35536
canonical_url
https://medium.com/@pmLearners/we-rewrote-8-of-the-system-in-rust-pager-volume-dropped-more-than-any-infrastructure-upgrade-fce168e35536
author_url
https://medium.com/@pmLearners
status
ok
fetched_at
2026-06-09 15:37:30