โ† Back to list

๐Ÿ’ฅ The Day a Single YAML Change Took Down Our Spring Boot Cluster

โ€œItโ€™s just a small config tweak. Whatโ€™s the worst that could happen?โ€ ย Famous last words.

Dolly ยท 2025-12-28 16:34 ยท 50 claps ยท 2.9 min read
#singles #yaml #took #spring-boot #cluster
Open on Medium โ†—

๐Ÿ’ฅ The Day a Single YAML Change Took Down Our Spring Boot Cluster

๐Ÿ’ฅ The Day a Single YAML Change Took Down Our Spring Boot Cluster

๐Ÿ’ฅ The Day a Single YAML Change Took Down Our Spring Boot Cluster

โ€œItโ€™s just a small config tweak. Whatโ€™s the worst that could happen?โ€ Famous last words.

This is the story of how one innocent line in application.yml brought down our entire Spring Boot cluster in under 7 minutes.

No code changes. No bad deploy. Just YAML.

๐Ÿงพ The system

We were running:

  • ~20 Spring Boot microservices
  • Kubernetes (EKS)
  • Java 21 + virtual threads
  • PostgreSQL + HikariCP
  • Kafka for async
  • ~25k RPS peak traffic

Solid. Stable. Or so we thought.

โœ๏ธ The change that started it all

A teammate wanted to โ€œoptimizeโ€ DB usage.

PR diff:

spring:
  datasource:
    hikari:
-     maximum-pool-size: 20
+     maximum-pool-size: 100

Thatโ€™s it.

The idea:

โ€œMore connections = more throughput.โ€

The PR sailed through. No alarms. No one expected trouble.

โฑ๏ธ Minute 0โ€“2: Looks fine

Pods rolled out. Health checks green. Dashboards normal.

We even congratulated ourselves:

โ€œNice, DB waits are gone!โ€

โฑ๏ธ Minute 3โ€“5: Latency creeps in

p95 latency: ๐Ÿ“ˆ 200 ms โ†’ 800 ms โ†’ 2 s

DB CPU: ๐Ÿ“ˆ 40% โ†’ 85%

Connections: ๐Ÿ“ˆ climbing fast.

But autoscaling hid it. More pods spun up.

We thought:

โ€œJust traffic.โ€

๐Ÿ’ฅ Minute 6โ€“7: Everything collapsed

Suddenly:

โŒ DB connections maxed out โŒ Queries timing out โŒ Pods stuck waiting for connections โŒ Kafka consumers lagging โŒ API gateway returning 503

Then: โ˜ ๏ธ Read replicas fell behind โ˜ ๏ธ Primary started rejecting new conns โ˜ ๏ธ Health checks failed โ†’ pods restarted

In less than a minute:

The whole cluster was down.

All from one line.

๐Ÿ˜ฑ What actually happened?

Before:

  • 20 pods ร— pool size 20 = 400 max connections

After:

  • 20 pods ร— pool size 100 = 2000 connections

Our Postgres max: โžก๏ธ 500 connections.

So we accidentally told:

20 services to DDOS our own database.

Each pod:

  • Opened more connections
  • Held them in transactions
  • Starved others
  • Caused timeouts
  • Triggered retries
  • Multiplied load

โžก๏ธ A self-inflicted denial of service.

๐Ÿ”„ The retry death spiral

To make it worse:

  • Timeouts triggered retries
  • Retries opened more connections
  • More connections slowed DB further
  • Slower DB caused more timeouts

A perfect feedback loop.

YAML + retries = outage amplifier.

๐Ÿ›‘ The fix (and the panic)

At minute 9:

  • We rolled back
  • Restarted pods
  • Manually killed idle connections
  • Paused consumers
  • Prayed

By minute 20: ๐ŸŸข Services slowly recovered.

Total impact:

  • ~15 minutes downtime
  • Thousands of failed requests
  • A very quiet Slack channel

๐Ÿง  The lessons that hurt the most

1๏ธโƒฃ Config scales multiplicatively

Per-pod limits ร— number of pods = real impact.

Never change:

  • DB pools
  • Thread pools
  • HTTP pools

Without calculating:

new_limit ร— max_pods

2๏ธโƒฃ Databases hate โ€œmoreโ€

DBs prefer: โœ”๏ธ fewer, well-used connections โŒ not thousands of idle ones

More connections โ‰  more throughput. Often it means more context switching and locks.

3๏ธโƒฃ YAML is production code

That one line had:

  • No tests
  • No canary
  • No guardrails

Yet it had:

More impact than most code changes.

Treat config as: โžก๏ธ code with blast radius.

4๏ธโƒฃ Autoscaling can hide disasters

HPA saw latency: ๐Ÿ“ˆ โ†’ added pods.

Each new pod: โžก๏ธ added 100 more DB connections.

Autoscaling didnโ€™t save us. It made it worse.

5๏ธโƒฃ We had no connection budget

We never defined:

โ€œHow many DB connections can this system afford?โ€

So YAML guessed for us.

Badly.

๐Ÿ›ก๏ธ What we changed after

Now, every service has:

spring.datasource.hikari.maximum-pool-size: 15

And cluster-wide:

  • DB connection budget documented
  • Per-service quotas
  • Alerts on:
  • active connections
  • wait time
  • Canary deploys for config
  • PR checklist for โ€œblast radiusโ€

We also added:

spring.datasource.hikari.connection-timeout: 2s

Fail fast > fail cluster.

๐Ÿ“ A simple safe formula

If DB max connections = 500 And worst-case pods = 20

Then:

pool_size โ‰ค 500 / 20 = 25

And keep headroom.

Never let YAML do this math for you.

โค๏ธ Final thought

The most dangerous changes in production arenโ€™t complex refactors. Theyโ€™re โ€œtiny safe tweaks.โ€

Because no one is scared enough to double-check them.

That day taught us:

A single YAML line can be more powerful than a thousand lines of Java.

Respect it.

If you enjoyed this story, share it with your team โ€” especially the one who says โ€œItโ€™s just config.โ€ ๐Ÿ˜„


๋ฉ”ํƒ€๋ฐ์ดํ„ฐ
post_id
b096ce6e9732
slug
the-day-a-single-yaml-change-took-down-our-spring-boot-cluster-b096ce6e9732
url
https://medium.com/@gangoladeepa/the-day-a-single-yaml-change-took-down-our-spring-boot-cluster-b096ce6e9732
canonical_url
https://medium.com/@gangoladeepa/the-day-a-single-yaml-change-took-down-our-spring-boot-cluster-b096ce6e9732
author_url
https://medium.com/@gangoladeepa
status
ok
fetched_at
2026-06-21 19:25:17