← Back to list

I Turned Off Autoscaling for 30 Days. Our AWS Bill Dropped 38% and Nothing Died.

We paid a 38% premium to let a robot react to our own deploys instead of our users.

The Speed Engineer · 2026-07-08 13:01 · 2 claps · 5.7 min read paywalled
#aws #kubernetes #devops #software-development #finops
Open on Medium ↗
Wiki topics: 🌐 · Web Development ☁️ · DevOps & Cloud

I Turned Off Autoscaling for 30 Days. Our AWS Bill Dropped 38% and Nothing Died.

We paid a 38% premium to let a robot react to our own deploys instead of our users.

Pinned capacity, literally — we stopped letting the dial move on its own and set it by hand. The bill fell and the building stayed warm.

Pinned capacity, literally — we stopped letting the dial move on its own and set it by hand. The bill fell and the building stayed warm.

Numbers are rounded and lightly anonymized; the ratios and timeline are real.

The $36K Insurance Policy Against a Threat That Never Showed

One number ended the argument: about $36,000 of compute last quarter ($12K a month), burned defending us from a traffic surge that never came.

Not “helped defend.” Burned. Every night, spinning up instances we didn’t need. It wasn’t reacting to our users. It was reacting to us.

For us, autoscaling had become a tax on not knowing our own load — and we’d paid it for years, because nobody had ever sat down and profiled our own baseline.

Nobody Audits the Autoscaler Because It Looks Like Best Practice

Our setup was nothing exotic: a predictable API service on EKS — diurnal traffic, a pile of cron jobs, CPU-based HPA. Not checkout, not payments, nothing where a delayed manual bump would put a customer-facing incident on the board. The monthly compute bill for this service ran around $32,000, and a slice of that was pure churn — roughly 40 scale-ups a day, each instance gone almost as fast as it arrived.

The autoscaler fired constantly, and nobody looked. Scaling on CPU is what you inherit — it’s on every “production-ready Kubernetes” checklist, it’s the default, and defaults are invisible. Invisible things don’t get audited; nobody opened the HPA history unless something was already on fire.

Here’s the trap with a cost you can’t see: you defend it without checking it. Every time finance asked about the bill, we said “it scales with traffic.” We never once confirmed that sentence was true.

“Elastic Equals Efficient” Is the Part That’s Wrong

The pitch was the one we parroted in every design review: capacity follows demand, you pay for what you use. We believed our own slides.

Pull the HPA history, though, and the embarrassing part jumps out: it had never measured demand. It watched one number we handed it, averageUtilization: 60 on CPU, and treated every bump as a customer knocking.

A deploy moves CPU. So does a GC pause. The cron job at the top of the hour? That too. They moved CPU — none of them moved request rate. An autoscaler doesn’t scale on demand; it scales on whatever metric you aimed it at, and we’d aimed ours at our own noise.

Three Weeks From Fear to a Flat Line

I wrote the original scaling policy. Copied the CPU target off a conference slide, shipped it, forgot it. It outlived two re-orgs.

The experiment was blunt: freeze autoscaling for 30 days, pin capacity at our p95 per-pod concurrency over a 30-day window plus 20% headroom, and watch. We left one exit — if p99 sat over the SLO for ten minutes, the HPA came back on. The rollback stayed in Git the whole time.

Week one, I didn’t trust my own math. I over-padded the headroom and refreshed the dashboards more times a day than I’ll admit. Latency never moved.

By week two, the sawtooth was gone. Twenty-four replicas, all day, every day. Same p99, same error rate. No SLO breach, no rollback, no customer ticket. Nothing died.

Elasticity wasn’t saving us from traffic. It was hiding a bad signal.

In week three, marketing ran a promo — a real spike, the exact event the autoscaler existed to catch. The pinned capacity absorbed it without me touching a thing, because the promo was smaller than the noise our own deploys had been throwing off.

It wasn’t spotless. One weekend a batch job collided with an organic traffic bump. The pager fired, I pulled up the dashboard, saw p99 drifting toward the SLO, bumped the floor from 24 to 30 by hand, and watched it settle. One page in 30 days — and if I’d been unreachable, it would have gotten uglier before anyone noticed. That’s the honest near-miss.

The One Line in Our HPA That Lied

Here’s what was driving the whole thing:

# hpa.yaml (before)
metrics:
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization
        averageUtilization: 60   # the line that lied — CPU load is not user demand

We pulled the HPA and pinned the deployment instead — removed from the active spec, rollback kept in Git:

# deployment.yaml (after) — capacity set by a human, reviewed monthly
spec:
  replicas: 24   # ceil(p95 concurrency ÷ per-pod capacity × 1.2)

We did weigh scaling on request rate or queue depth instead of pulling it. But the load was predictable enough that a pinned floor was cheaper to run — and far easier to audit — than tuning a new signal.

The waste showed up the moment we queried the scale-up events:

# Correlate HPA scale-up events with deploy + cron markers
changes(kube_horizontalpodautoscaler_status_desired_replicas[10m]) > 0

Overlaid on our deploy timestamps and the top-of-the-hour batch cron, the scale-ups tracked deploys, GC pauses, and cron — almost never request rate.

The replacement fit on one page: 30 days of per-pod concurrency, take the p95, add 20%, round up to a replica count, and put a name next to it — the service owner, not the platform team — for the monthly review.

Four Questions to Ask Before You Trust an Autoscaler

Don’t copy our number — copy the audit.

First: what signal actually fires it? Not “traffic” — the literal metric. Go read the YAML. If it says CPU, you’re scaling on CPU, and for us CPU was a liar.

Second: how often did it fire last month, and against what? Overlay the scale-up events on your deploys and cron. If the peaks line up with internal activity, elasticity is masking noise, not absorbing demand.

Third: what did each firing cost? Instances times minutes times rate. Put a dollar figure on it — ours was $36,000 of pure noise. An abstract “it scales” turns into a line item the second you multiply it out.

Fourth, the one people skip: what breaks if it never fires again? If the honest answer is “nothing, as long as we sized the floor right,” you weren’t buying elasticity. You were buying insurance against a number you never measured.

What Static Capacity Actually Cost Us

Monthly compute for this service fell from about $32,000 to about $20,000 — that’s the 38%. Across the quarter, roughly $96K down to $60K.

Before (HPA on CPU) After (pinned) Instances sawtooth, ~40 scale-ups/day flat at 24, one manual bump to 30 Monthly compute ~$32K ~$20K p99 / errors within SLO unchanged, no breach

It wasn’t free. Capacity planning is a human job again: the service owner runs the monthly review, pulls the p95, re-pins the number. The on-call runbook changed too — “the autoscaler will handle it” became “check the pinned floor, and if p99 sits above the SLO, bump replicas ~20% and page the capacity owner.” And I lost one Saturday to a spike that elastic capacity might have quietly swallowed.

That’s what we actually bought: cheaper boxes, and one person who now has to care, every month, whether those boxes still fit.

Don’t Copy This If Your Load Is Spiky

Turn your autoscaler off only if three things are true. Your load is predictable and diurnal — it breathes in and out on a rhythm you can chart. You’ve done the headroom math, not eyeballed it. And a specific person, by name, owns a monthly capacity review, or the whole thing rots.

If your traffic is genuinely bursty — real customer spikes several times baseline, unpredictable — keep the autoscaler. Match your scrutiny to your blast radius. Static capacity on spiky load is just an outage with a countdown.

And the point isn’t “never autoscale.” It’s autoscale on the right signal, with an owner and a cost review — or stop pretending it’s free.

This doesn’t end clean, either. The bill’s already creeping back as teams stand up new services, each one shipping the same copied CPU policy I wrote years ago.

The signal owns the bill — unless someone owns the signal.

Autoscaling was never a knob you set once. It’s a control loop, and it needs a human in it.

Enjoyed the read? Let’s stay connected!

  • 🚀 Follow The Speed Engineer for more Rust, Go and high-performance engineering stories.
  • 💡 Like this article? Follow for daily speed-engineering benchmarks and tactics.
  • ⚡ Stay ahead in Rust and Go — follow for a fresh article every morning & night.

Your support means the world and helps me create more content you’ll love. ❤️


메타데이터
post_id
bbd6ef050f16
slug
i-turned-off-autoscaling-for-30-days-our-aws-bill-dropped-38-and-nothing-died-bbd6ef050f16
url
https://medium.com/@speed_enginner/i-turned-off-autoscaling-for-30-days-our-aws-bill-dropped-38-and-nothing-died-bbd6ef050f16
canonical_url
https://medium.com/@speed_enginner/i-turned-off-autoscaling-for-30-days-our-aws-bill-dropped-38-and-nothing-died-bbd6ef050f16
author_url
https://medium.com/@speed_enginner
status
ok
fetched_at
2026-07-09 13:13:48