๐ฅ The Day a Single YAML Change Took Down Our Spring Boot Cluster
โItโs just a small config tweak. Whatโs the worst that could happen?โ ย Famous last words.

๐ฅ The Day a Single YAML Change Took Down Our Spring Boot Cluster
๐ฅ The Day a Single YAML Change Took Down Our Spring Boot Cluster
โItโs just a small config tweak. Whatโs the worst that could happen?โ Famous last words.
This is the story of how one innocent line in application.yml brought down our entire Spring Boot cluster in under 7 minutes.
No code changes. No bad deploy. Just YAML.
๐งพ The system
We were running:
- ~20 Spring Boot microservices
- Kubernetes (EKS)
- Java 21 + virtual threads
- PostgreSQL + HikariCP
- Kafka for async
- ~25k RPS peak traffic
Solid. Stable. Or so we thought.
โ๏ธ The change that started it all
A teammate wanted to โoptimizeโ DB usage.
PR diff:
spring:
datasource:
hikari:
- maximum-pool-size: 20
+ maximum-pool-size: 100
Thatโs it.
The idea:
โMore connections = more throughput.โ
The PR sailed through. No alarms. No one expected trouble.
โฑ๏ธ Minute 0โ2: Looks fine
Pods rolled out. Health checks green. Dashboards normal.
We even congratulated ourselves:
โNice, DB waits are gone!โ
โฑ๏ธ Minute 3โ5: Latency creeps in
p95 latency: ๐ 200 ms โ 800 ms โ 2 s
DB CPU: ๐ 40% โ 85%
Connections: ๐ climbing fast.
But autoscaling hid it. More pods spun up.
We thought:
โJust traffic.โ
๐ฅ Minute 6โ7: Everything collapsed
Suddenly:
โ DB connections maxed out โ Queries timing out โ Pods stuck waiting for connections โ Kafka consumers lagging โ API gateway returning 503
Then: โ ๏ธ Read replicas fell behind โ ๏ธ Primary started rejecting new conns โ ๏ธ Health checks failed โ pods restarted
In less than a minute:
The whole cluster was down.
All from one line.
๐ฑ What actually happened?
Before:
- 20 pods ร pool size 20 = 400 max connections
After:
- 20 pods ร pool size 100 = 2000 connections
Our Postgres max: โก๏ธ 500 connections.
So we accidentally told:
20 services to DDOS our own database.
Each pod:
- Opened more connections
- Held them in transactions
- Starved others
- Caused timeouts
- Triggered retries
- Multiplied load
โก๏ธ A self-inflicted denial of service.
๐ The retry death spiral
To make it worse:
- Timeouts triggered retries
- Retries opened more connections
- More connections slowed DB further
- Slower DB caused more timeouts
A perfect feedback loop.
YAML + retries = outage amplifier.
๐ The fix (and the panic)
At minute 9:
- We rolled back
- Restarted pods
- Manually killed idle connections
- Paused consumers
- Prayed
By minute 20: ๐ข Services slowly recovered.
Total impact:
- ~15 minutes downtime
- Thousands of failed requests
- A very quiet Slack channel
๐ง The lessons that hurt the most
1๏ธโฃ Config scales multiplicatively
Per-pod limits ร number of pods = real impact.
Never change:
- DB pools
- Thread pools
- HTTP pools
Without calculating:
new_limit ร max_pods
2๏ธโฃ Databases hate โmoreโ
DBs prefer: โ๏ธ fewer, well-used connections โ not thousands of idle ones
More connections โ more throughput. Often it means more context switching and locks.
3๏ธโฃ YAML is production code
That one line had:
- No tests
- No canary
- No guardrails
Yet it had:
More impact than most code changes.
Treat config as: โก๏ธ code with blast radius.
4๏ธโฃ Autoscaling can hide disasters
HPA saw latency: ๐ โ added pods.
Each new pod: โก๏ธ added 100 more DB connections.
Autoscaling didnโt save us. It made it worse.
5๏ธโฃ We had no connection budget
We never defined:
โHow many DB connections can this system afford?โ
So YAML guessed for us.
Badly.
๐ก๏ธ What we changed after
Now, every service has:
spring.datasource.hikari.maximum-pool-size: 15
And cluster-wide:
- DB connection budget documented
- Per-service quotas
- Alerts on:
- active connections
- wait time
- Canary deploys for config
- PR checklist for โblast radiusโ
We also added:
spring.datasource.hikari.connection-timeout: 2s
Fail fast > fail cluster.
๐ A simple safe formula
If DB max connections = 500
And worst-case pods = 20
Then:
pool_size โค 500 / 20 = 25
And keep headroom.
Never let YAML do this math for you.
โค๏ธ Final thought
The most dangerous changes in production arenโt complex refactors. Theyโre โtiny safe tweaks.โ
Because no one is scared enough to double-check them.
That day taught us:
A single YAML line can be more powerful than a thousand lines of Java.
Respect it.
If you enjoyed this story, share it with your team โ especially the one who says โItโs just config.โ ๐
๋ฉํ๋ฐ์ดํฐ
- post_id
- b096ce6e9732
- slug
- the-day-a-single-yaml-change-took-down-our-spring-boot-cluster-b096ce6e9732
- url
- https://medium.com/@gangoladeepa/the-day-a-single-yaml-change-took-down-our-spring-boot-cluster-b096ce6e9732
- canonical_url
- https://medium.com/@gangoladeepa/the-day-a-single-yaml-change-took-down-our-spring-boot-cluster-b096ce6e9732
- author_url
- https://medium.com/@gangoladeepa
- status
- ok
- fetched_at
- 2026-06-21 19:25:17