We Didn’t Migrate Systems. We Migrated Assumptions
What moving an invoicing platform off Heroku onto EKS actually cost, and the three failures nobody puts in the slide deck
We Didn’t Migrate Systems. We Migrated Assumptions
What moving an invoicing platform off Heroku onto EKS actually cost, and the three failures nobody puts in the slide deck

It was a Tuesday afternoon when both environments went down at once.
We had pushed a routine deploy to the new EKS cluster. Thirty seconds later the connection pool was exhausted. Sixty seconds later the new pods were ignoring shutdown signals and leaking connections. Two minutes later the leak had reached across the network and drained the connection limit on the old Heroku Postgres too. The platform we were migrating from and the platform we were migrating to were both on the floor, at the same time, because of one line in a Dockerfile.
That afternoon is the reason I’m giving this talk at Open Source Summit North America 2026 in Minneapolis on Monday. Not because the migration succeeded, although it did. Because of how badly it went before it did, and because almost nothing I read beforehand prepared me for the specific ways it would break.
This is the long version of that talk.
The starting point
The product was a fast-growing invoicing SaaS. Around 2 million active small business merchants, roughly 33 million invoices a year, enterprise customers with SLAs written into contracts. The architecture was already 47 Node.js microservices on Heroku, SQS for events, Redis for sessions. The engineering team was 10 people. Two of us owned the platform.
The talk is called “Monolithic to Cloud Native,” and people always push back on that title, fairly. The services were already micro. So what was the monolith? The platform was. Every one of those 47 services routed its assumptions through one PaaS that quietly made decisions on their behalf, and those decisions were precisely the ones that detonated the moment we left.
By early 2026 four things were going wrong at the same time and all of them were getting worse. API latency had settled at 700ms at the 99th percentile, and there was no lever left to pull because we had hit the Heroku dyno scaling ceiling. The deploy pipeline took 45 minutes or more, against enterprise SLAs we kept breaching. We had no container-level observability, so most of our debugging was educated guessing. And the monthly bill had crossed the line where it cost more than it returned.
We sat with the options honestly. Stay on Heroku and accept the ceiling. Rewrite for serverless and pay the rewrite. Move to raw AWS VMs for cost relief but no velocity. Or move to EKS, which we knew going in was the highest-risk, highest-ceiling path on the board. We chose EKS with our eyes open about which one of those it was.
Three failures before it worked
The first failure was invisible. The PDF generation service went from 800ms to 9 seconds, and every dashboard we had showed CPU sitting at a comfortable 35%. It took days to understand that the Linux CFS scheduler hands out CPU in 100ms slices, and a 500m limit means you get 50ms of every 100ms. Node.js spawns a handful of worker threads, garbage collection runs on its own, and six threads were fighting over that 50ms window. Average CPU looked low because the process spent most of its life throttled rather than running. Heroku had been letting dynos burst, so we had never seen this. A CPU limit, it turns out, is not a number. It is a scheduling rule, and the rule was strangling us in plain sight.
The second failure was multiplicative. Heroku’s resolver was configured with one DNS search behavior. The EKS default is another, ndots:5, which makes the resolver walk a list of search domains before trying a hostname as written. A name like api.stripe.com quietly became roughly 10 DNS lookups. We made 150,000 of those calls a day. That turned into 1.5 million queries, and across every external integration, around 12 million unnecessary DNS queries a day. CoreDNS was the thing buckling, and the cause was a default we never chose and never saw.
The third failure was the Tuesday afternoon. A Dockerfile that started the app in shell form meant the process running as PID 1 was a shell, and shells swallow the shutdown signal instead of passing it on. The Node process never learned it was supposed to drain and exit. It just held its connections until the pool overflowed, and then it took the shared database down with it, on both sides of the migration.
What unsettled me was not any single bug. It was the question underneath all three. Why did none of our dashboards catch any of this? Because the CFS default was invisible, the DNS amplification was multiplicative rather than additive, and PID 1 betrayed us in a place nothing was watching. We had not been migrating systems. We had been migrating assumptions, and the assumptions were the part that failed.
The architecture we landed on

People ask why we didn’t just script AWS directly. The honest answer is that four decisions cost less to make once at the platform layer than to carry forever in every team. Every one of them is open source, and every one of them is the reason a two-person team could run this.
We rejected DNS-based routing and load balancer weighting and chose Istio for traffic shifting. Canaries moved in steps, 5%, 25%, 50%, 100%, and a rollback was a single config change that took seconds with no redeploy and no DNS propagation. Mutual TLS arrived for free with the mesh. Before we moved a single byte we spent two weeks collecting a baseline with Prometheus, Thanos, and Grafana, watching Heroku and EKS on the same panels, because you cannot migrate what you cannot measure. Infrastructure changes went through Atlantis, so a Terraform plan landed in a pull request comment and nobody ran apply from a laptop at 2am ever again. And deploys became git commits through Flux, with drift detection quietly correcting the manual changes everyone swears they didn’t make.
The database cutover used a dual write into RDS with checksum-validated replication running continuously underneath. When we finally flipped it, the cutover was anticlimactic. That was exactly what we had been working toward.
The results

The headline numbers, from Heroku to EKS:
- API latency p99: from 700ms to 70ms, down 90%
- Deploy time: from 45 minutes to 4 minutes, down 91%
- Monthly incidents: from 12 to 2, down 83%
- Deploy frequency: from twice a week to 15 a day, up 50x
- Monthly cost: more than 60% lower, from right-sizing, spot, and Karpenter
I don’t trust a 90% number with no story behind it, so here is where the 630ms actually went. Roughly 250ms was routing variance, Istio’s least-connection routing against Heroku’s effectively random routing. Around 160ms was network topology, pod-to-pod inside the VPC instead of a public path with TLS renegotiation. About 125ms was resource isolation, with CFS throttling falling from 65% of scheduling periods to under 2%. The last 95ms was connection pooling through PgBouncer in transaction mode.
And the line that belongs in every migration writeup and almost never appears: this absorbed two platform engineers full-time for five months, plus about 30% of eight application engineers’ time. Nothing here was free, and pretending otherwise is how teams talk themselves into migrations they aren’t staffed for.
The quieter win was developer experience. Heroku’s superpower was a single git push, and we were never going to beat that, so we got close with an internal developer portal built on Backstage. A new service scaffolds in about five minutes. Kubernetes complexity stayed our problem, not the developers’ problem. That is how a 10-person team grew to 100 on a platform that two of us maintained, with the service count holding at 47 the entire time because the platform absorbed the growth instead of the codebase.
Lessons learned
Every platform hides a different class of failure, and the failures it hides are the ones that become critical the moment you leave. The defaults you never had to think about are the defaults that will take you down.
Instrument before you migrate, not during. The two weeks of baseline felt like a delay at the time. It was the single highest-leverage thing we did, because every later number had something to be compared against.
Reversible beats fast. The migration that worked was the one where every step could be undone in seconds, and the database cutover was boring on purpose. Boring is the goal.
And the part I keep coming back to. A two-person platform team in Ontario ran the same infrastructure stack as companies a hundred times our size. That is only possible because thousands of contributors built the tools we stood on, which is why we put 22 merged pull requests back across 14 of those projects this year. Open source is the equalizer. It is the reason a small team at a mid-market company can run infrastructure that used to require a department.
If you’re at Open Source Summit North America 2026, the talk is Monday, May 18, at 5:25pm CDT in Room 200F at the Minneapolis Convention Center. I’ll stay after for the parts that don’t fit in 25 minutes. There are a lot of them. Slides and the full contribution list are at phonotech.ca/ossna26.
Have thoughts on platform migrations, or war stories of your own? Let me know in the comments.
메타데이터
- post_id
- 50eda2c3d9bc
- slug
- we-didnt-migrate-systems-we-migrated-assumptions-50eda2c3d9bc
- url
- https://medium.com/@mateenanjum/we-didnt-migrate-systems-we-migrated-assumptions-50eda2c3d9bc
- canonical_url
- https://medium.com/@mateenanjum/we-didnt-migrate-systems-we-migrated-assumptions-50eda2c3d9bc
- author_url
- https://medium.com/@mateenanjum
- status
- ok
- fetched_at
- 2026-07-10 14:51:46