← Back to list

My Experience at SRECon EMEA ‘25

Structure of this Post

Rob Durst in Spring Health Engineering · 2025-11-17 20:34 · 1 claps · 11.5 min read
#site-reliability-engineer #sre #srecon #software-engineering #site-reliability
Open on Medium ↗

My Experience at SRECon EMEA ‘25

I’m a speaker!

I’m a speaker!

Structure of this Post

This post is an overview of my experience at SRECon EMEA ’25. I realize not everyone will be interested in the verbosity of the full narrative. So, I have put learnings in code sections:

==== Learning ====

Some Learning.

I won’t judge if you just skim to the learnings 😁

Overview

This past month I attended SRECon EMEA ’25 in Dublin, Ireland where I had the wonderful opportunity to speak on behalf of Spring Health about our journey to SLO Readiness (*video*). This was a trifecta of firsts for me:

  1. first time in Dublin, Ireland (fun fact I grew up most of my life in Dublin, CA)
  2. first time attending SRECon in person (binged a bunch of past conference videos, but never been)
  3. first time giving a conference talk

While (1) was very cool and (2) an incredible learning opportunity, (3) was special as ever since co-organizing the SF Cryptocurrency Developer Meetup in SF, speaking at a conference has been on my bucket list.

Before going further, I want to acknowledge Spring Health’s financial support to attend this conference!

Some General Thoughts

I am not the most outgoing conference goer. Typically I don’t feel comfortable initiating conversations and this compounded with being the sole member of Spring Health to attend (and it felt like maybe even the health care industry).

==== Learning ====

Healthcare representation seemed limited, which surprised me; reliability is
such a critical part of care delivery. Really too bad in my opinion as our
notion of care blocking being the framework for understanding incident impact
is powerful. Raised this in the feedback for organizers, but much like the
"Reliability in Finance"  track, a "Reliability in Healthcare" track would
be awesome!

There was likely more to be desired in terms of making connections with my peers in industry. However, many learnings were still had, connections made, and giving a talk was a great experience for myself (and surfacing some of the awesomeness that is Spring’s engineering organization).

Ok, Dublin was Pretty Great

First off, minutes after exiting the terminal, I came across numerous specialty coffee roasters. Incredible. My favorite was 3fe in the Financial District. Went there nearly every morning before the conference (and yes, even though there was free coffee at the event…).

3fe Coffee Roasters

3fe Coffee Roasters

Also, and I’m not a drinker anymore, but the open Guinness bar during the reception event was kind of epic. And finally the venue was beautiful, right on the River Liffey. Here is a pic I snapped (during a late night session walking and mentally practicing my presentation 😅).

Dublin Convention Centre (green)

Dublin Convention Centre (green)

Oh ok, bonus pic. Loved the Irish breakfast.

Irish Breakfast

Irish Breakfast

Overall, most of my time was spent at the conference however I did get a small taste of Dublin. Also took a post conference detour to London where I saw England beat Wales at Wembley. Feel free to ask me about that offline — good time but not the topic of this post.

My Talk: “Run, Walk, Crawl, or How We Failed Our Way to SLO Readiness”

Backstory

At Spring we’ve come a long way in the three and a half years I’ve been here. Motivated to see if I could come up with a topic for a talk, it really felt like there was a story to tell. I did a lot of reflection, chatted with numerous teammates, and ultimately came up with a story describing the evolution of our SRE Culture (which I first presented internally at our company onsite). Condensing down that talk to a shorter, 15 minute version and focusing on a specific thing SRE’s especially love — SLOs — I submitted my proposal ~45 minutes before the deadline (a large part of which was my wife saying “eh, what’s the worst that can happen?!?” She’s great, if you love the talk, thank her… it almost didn’t take place!).

Abstract

Per the short description on the conference website:

Scaling a site-reliability culture from the ground up at a hyper-growth, resource-constrained startup is a uniquely challenging endeavor: plenty of playbooks to scale SRE teams exist, yet every startup’s socio-technical reality is its own puzzle. And while “reliability is the most important feature” rings true, tight deadlines and shifting priorities often sideline proactive reliability initiatives. Thus, since these cycles are a precious commodity, ensuring their success is paramount.

By retracing our SLO adoption journey, highlighting failures, missteps, and near wins en route to our eventual breakthrough, we uncover an effective litmus test for gauging readiness (or recognizing when a team isn’t quite there yet).

Today this readiness framework guides how we assess the timing of reliability investments at Spring Health. It also serves as a practical tool for teams in fast-growing engineering orgs still early in their reliability journey, especially those navigating similar constraints.

My Experience

The Talk

First off, a lot of respect was gained for conference speakers. My talk was only fifteen minutes, but the time spent refining the message, outlining the story, putting together slides, rehearsing these slides, etc. was fairly substantial. Yet when all things were said and done and I was up at the mic in front of ~50–100 SREs (no clue the capacity of the room), it went well! The jokes landed and the content was well received: a few folks approached me afterwards with “great talk” and “we can totally relate, that was validating.” However I must say the Q&A did not go as well as I’d hoped; I had spent so much time preparing for the talk, I was not super well prepared for the Q&A. Sure it went fine, but some of the questions were predictable and I could probably have prepared better answers ahead of time.

==== Learnings ====

If it says there is a Q&A, there really is a Q&A… so prepare.

We Can Do It

I discuss this theme in more detail below, but really (a) having this talk be accepted for what it was and (b) having folks acknowledge its usefulness was validating not only just for me as the speaker, but as an engineer at Spring. It’s easy to sometimes write off the work we do and not give ourselves the full credit we deserve. But we’re doing some really cool things. We can hold our place in industry. I hope others of similarly sized orgs see this talk and consider sharing more about the wonderful and interesting work they’re doing!

Artifacts

Slides: https://www.usenix.org/conference/srecon25emea/presentation/durst

Video:

[embed]My talk!

And a Bonus Talk — Mama Look, I Compiled SLOs!

A few folks ended up dropping out of the Lightning Talks, so I eagerly signed up (with some regret later realizing quite quickly after the initial excitement that I now had two things to be nervous about). For this I spoke outside my “official Spring Health” capacity about the work on what has now become a full blown compiler for “generating reliability artifacts from system specifications” which I am hacking on in my spare time but leveraging at Spring (and also giving a talk about at the first ever Gleam conference this winter!).

Final slide of my talk.

Final slide of my talk.

Lightning talks are fun and with much lower stakes… however if I ever do one of these again I’d pick less of a passion project and optimize more for entertainment value (believe it or not the normal folk is not overly passionate about compilers).

Discussion Tracks

On the second day, I came to a realization: I realized the discussion tracks would help me overcome the difficulty of being more of an introvert and furthermore, all conference talks are recorded, discussions are not, so at the very least I’d be part of something unique to being a participant at the event. I attended two discussions:

  1. Charity Major’s “On Building Systems Where Normal Engineers Can Do Great Work” inspired by her blog post In Praise of Normal Engineers
  2. Alex Hidalgo’s “Your Observability Is Expensive (and So Are Your Feelings)

On Building Systems Where Normal Engineers Can Do Great Work

Charity’s discussion was well attended with many folks giving their unique perspective on the topic. Even though it was a full hour and a half, it went by pretty fast and the conversation flowed.

==== Learnings ====

* Good orgs mint excellent engineers
* Typically heroic engineers are just folks with skill + more context. 
  If we can scale context, we can scale engineering.
    * “biggest carrying costs are cognitive load and so only smart people
      can ship if feedback loop too big”
* One way to fix the context issue is with abstractions (reduced cognitive
  load, reduced tedium). However, too much abstraction can put “blinders on”
  engineers such that they don’t really understand what’s happening.
* Have to give people tooling if have expectations

Your Observability Is Expensive (and So Are Your Feelings)

Alex Hidalgo’s session wasn’t quite as full, but it did foster a good discussion on the cost of an observability platform. Since at SRECon vendors are well represented, the discussion was not centered on vendor specific gripes, but the cost as a whole. Furthermore, I have to say, it felt like a real privilege being a participant in these discussion rooms as these conferences still draw the OGs of industry: at just this discussion track (in a room of maybe 30 engineers) was Alex Hidalgo, Todd Underwood, and Liz Fong Jones to name a few.

==== Learnings ====

* On a discussion of paging and noise, critical point that working hours
  are not the same across the board (i.e. parents vs. party-goers).
* [My point] when held wrong, we incur an educational cost.
* My notes here leave a lot to be desired, but overall framing the cost
  of observability as all the costs beyond just $$$ is an interesting
  conversation we might consider replicating here at Spring.

AI

At first I thought “oooh, I’ll go to a bunch of AI talks and come back with a ton of things to implement at Spring”. However, after the first couple talks, I opted for more discussion tracks, so will need to do some session catchup when videos are released.

Nevertheless, I still went to some talks in this domain.

Why Risk Management Requires Taking Risks: A Practical Guide to Getting SRE Teams AI-Ready

I went to a talk by Nvidia titled “Why Risk Management Requires Taking Risks: A Practical Guide to Getting SRE Teams AI-Ready”. They mostly talked about how to interject AI into processes. Their approach around intentionally engineering context to set AI up for success was interesting (and then went deep into their Bayesian Network they’d developed for this). Furthermore their highlighting of recommendations for a successful rollout were pretty straightforward but practical: ideas like start small, bridge skill gaps, etc.

==== Learnings ====
Leveraging AI within processes in a read-only capacity that augments engineers
plays well with the discussion above around empowering the “normal” engineer.
Worth exploring this more.

MLOps

MLOps 2025: A Journey into the Past and the Future

In my opinion, something hard to miss at this year’s conference was the recurring theme of MLOps. MLOps, while not new (as wonderfully explained in the opening plenary session “MLOps 2025: A Journey into the Past and the Future”), is starting to really become more widely adopted as a practice and even a role as more companies ship more AI/ML services; this talk specifically addressed (tongue in cheek) LLMOps and stressed it is just “renaming the same foundations.”

==== Learnings ====

* MLOps borne from 2015 paper “Hidden Technical Debt in Machine Learning
  Systems” 
* MLOps will be commonplace. Not new. Same foundations will power: AIOps,
  LLMOps, etc.
    * However, now focus on inference centric vs. originally training
      centric
* Huge gap in MLOps monitoring practices, may not be common place until 2030
    * Survey says 50% of folks have little to no monitoring‼️
* Predicting the next abstraction of SRE will come about around 2033 and
  well likely need new programming languages to handle new abstractions
* In terms of current AI landscape, down the SDLC funnel, maturity from code
  to test to operate to monitor and debug is less and less and less mature.

SRE for AI and AI for SRE

Another fascinating talk within this realm was Todd Underwood’s “SRE for AI and AI for SRE”. He dove deep into a recent Anthropic incident, highlighting more around AI and SRE. Specifically of note, he discussed the process for debugging these types of complex issues, noting that it is exceptionally hard due to the limited insight into what the user experienced and the non-deterministic nature of LLMs.

==== Learnings ====

* Read Reliable Machine Learning
* Interesting to get ahead on thinking hard about signals we’d wish we’d have
  when debugging incidents
    * Furthermore, figuring out ways to incentivize users to give us feedback
    * Can we attach attribute to “thumbs up, thumbs down” and treat these as
      wide events?

Relatable — We’re All Different but Also All The Same

Really, one of my biggest takeaways from this experience was validation 😁

SRE is a Role

In the #hallway slack channel for the conference, following my talk, I was curious what SRE looked like at other engineering organizations.

Many thanks to the wonderful folks who responded to that thread! In the first 5 years and change of my career since graduating from college, I’ve only actually donned the title “SRE” for a year and a half (and even technically speaking, within Spring’s title system it is DevOps). However, the scope of work I’ve taken on (including time at DocuSign) has included: software load balancing, load testing, rightsizing containers, incident management, lots of observability work etc.

And thus, I believe I have a strong case to argue I’ve really been within the SRE role for a majority of my professional career.

==== Learnings ====
While the role versus title discrepancy may seem pedantic, for folks who
feel imposter syndrome (especially as many at SRECon would agree SRE is a
second job type role), this discrepancy is a great way to frame things and
a wonderful validation of one’s abilities.

We’re Actually All the Same?

This theme of validation manifested in other ways as well. Throughout the conference, a theme emerged: while we work at companies of different sizes, within different domains, with different cultures, and different maturities, we’re all really solving similar problems.

Some specific examples:

Run, Walk, Crawl, or How We Failed Our Way to SLO Readiness

[Company]: Spring Health (me)

[Similarities]: received feedback from numerous folks that they were going through a similar journey and it was validating to hear my perspective.

CPU Utilization: The Hidden Cost of Running Hot

[Company]: Github

[Similarities]: discussed how they do performance testing of different hardware configs. We’ve done similar tuning ourselves. Also shoutout to another Ruby shop. To do this they had a whole team focused on this as the wins were large enough to justify that. Obviously we can’t do that yet, but easy to reason about how, if we continued scaling, even just dedicating a sprint to this would make sense.

Taming the Cost of Telemetry: How Riot Games Reined In Observability Costs

[Company]: Riot Games

[Similarities]: they have a team focused on owning, understanding, and evangelizing how to contain observability costs. Discussed the issues with cardinality (Rob vs. cardinality with custom metrics is an issue we have…). Didn’t see all this talk due to answering async questions from my own but look forward to it when it’s out.

Training New Incident Commanders: Pokemon Style!

[Company]: Datadog

**[Similarities]:** a discussion on how they developed and leveraged an incident commander training scenario. First off, this is awesome and they said they’d be happy to share — again, as a smaller org we’ve talked about this but never prioritized it. Furthermore all their K8s clusters are named after Pokemon. Nice to know we’re not the only crazy ones who like to name things after Pokemon…

==== Learnings ====
While folks at places like Google, Apple, Amazon, etc. are doing some
seriously crazy and impressive things that’ll never be relevant here,
the average joe SRE at places like Datadog, Github, etc. face many of the
same challenges and obstacles we do, just at companies of different sizes,
within different domains, with different cultures, and at different maturity
levels.

My Five Action Items

  1. Really dive in to MLOps. As a concrete action item: read Reliable Machine Learning. Maybe start a book club.
  2. Continue our SLO journey. SREs across the board swear by these.
  3. Signup for SRECon ’26 and convince teammates to go.
  4. Pontificate more on the idea of empowering “normal” engineers.
  5. Reduce educational cost of leveraging observability at Spring.

If you find this type of work interesting, check out our open positions. We’re hiring for several different engineering roles, and we’d love to speak with you.


메타데이터
post_id
e09db8b3e2b0
slug
my-experience-at-srecon-emea-25-e09db8b3e2b0
url
https://medium.com/spring-health-engineering/my-experience-at-srecon-emea-25-e09db8b3e2b0
canonical_url
https://medium.com/spring-health-engineering/my-experience-at-srecon-emea-25-e09db8b3e2b0
author_url
https://medium.com/@rdurst_46653
status
ok
fetched_at
2026-07-15 07:10:03