Incident Management: Farming in a Crisis
While we recorded a recent episode of This is Fine!, Beth Adele Long managed to plant a seed in my brain. In the middle of talking about…
Incident Management: Farming in a Crisis
While we recorded a recent episode of This is Fine!, Beth Adele Long managed to plant a seed in my brain. In the middle of talking about the paper she co-authored with the late Dr. Richard Cook, she pointed out that the “software factory” model (and really, Taylorism in general) is flawed. She suggested another analogy: software is grown in the soil of your working conditions.

At the time I sort of nodded and agreed, and it made a ton of sense, but didn’t think much of it. But, over the next few weeks, the seed sprouted, and started putting down roots. So I kept watering it. Meanwhile, I noticed the severe juxtaposition between this idea, and how we tend to build our software farms. We build an awful lot of factories:
CI/CD, we need a pipeline.
Product Life Cycle, we need a set of linear progressions from idea to completion.
Incident Management, we need phases.
Sometimes this makes sense. The repeatable, observable nature of a factory production line is extremely attractive. Did we do X yet? Did Y go well? How many defects had to be rejected? In fact, even in a farm, we have a series of steps: till the soil, fertilize, plant, water, weed, harvest, etc. The models are not totally incompatible.
But when we zoom out to the whole system, we start to see that the factory model misses a giant portion of what makes teams successful. It does nothing to address the quality of soil in which they’re growing software.
Who among us has not seen our organization invest in a great PLC process, and an adequate CI/CD system(nobody ever has a great one… for long), only to find that still, nothing gets done. And when we look at why, we see that the conditions for creating software are simply not present anymore. Perhaps we don’t have adequate mentorship, or leadership, and people are not gaining the skills necessary. Or perhaps we’ve constrained creativity, and burdened everyone with maintaining these linear processes.
If we aren’t checking on the conditions, and building processes that renew our soil, we shouldn’t be surprised when yields are disappointing.
So that’s all great, but, how does this relate to a crisis? Well, I think we often do this during incidents too. Incidents are expected to move, from detected, to investigated, to mitigated, to resolved. At each phase we have all sorts of activities that we deem need to happen. Status page update? Notify stakeholders? Submit break-glass approvals? Document duration? Assess impact, etc. etc. etc.
Making sure those happen is not a bad idea. And having dedicated folks and systems to ensure they are not a huge burden is worth an investment. Your incident management process is not a waste.
However, this is where Beth’s seed started to grow beyond the little plot in my head where “software development” lives. It twisted its way into the moment of crisis in which we need to figure things out. It occurred to me that the best incident outcomes are the ones where the conditions are right for the right kind of creativity to blossom in a moment of crisis.
This came to a head for me during a series of recent incidents that had very, very different shapes.
The first one lasted for many days, with intermittent but persistent failures affecting a narrow band of users. The beginning was normal-feeling. Find what changed, triangulate on the point of impact, reproduce the error, etc.
But none of that worked. The problem just wasn’t yielding its secrets. Meanwhile, conditions worsened for incident responders. Handovers became expensive as the context grew massive. More customers were upset, more parties got invovled, and detection remained very difficult. Pressure ramped up, coordination costs skyrocketed. Creativity yielded to incident management and coordination activity. New party has to ask all the same questions. Did you try this? Is it that? Have you tried that again since then? And whether you want it to or not, as the time goes on, there is a deep organizational need to blame somebody and at some point peoples patience runs out. The hypothesis crop yield was extremely meager. The soil was completely depleted.

Only when we wrestled command back from this paradigm did we manage to once again create the conditions for creativity to continue. We started to renew the soil as best we could. We discarded stale context and re-visisted now weeks-old assumptions. We separated the communication and user support function from the technical response. We reminded responders that while everyone hopes we can figure the whole thing out, right now we need to mitigate it by any (safe) means. We did our best to slow things down for the technical responders, to get them to focus on likely hypotheses, and to get help where we could, and blockers out of their way.
Finally, with the space and constraints set appropriately, the crop yielded fruitful hypoteheses. Creative engineers found the boogieman in the system, and managed to mitigate the issue.
Then almost immediately as that long, slow, frustrating incident resolved, a new one started.
This time the problem was acute and well-understood, but the constraints meant that the normal growing season for solutions would not be appropriate. This wasn’t a thing that could wait days. In fact, in some cases customers had just a handful of hours before consequenes would become untenable for them.
We didn’t panic. We ran the incident process. We didn’t let the soil get depleted with pressure this time. Immediately we separated the tech response from the communication and user support function. We got help that we needed, and we removed blockers. We planted anything and everything that we knew could grow in the soil we had.
Creative solutions sprouted quickly. Some of them weren’t right, and some were. In this extreme time-limited, high-pressure situation, we managed to create conditions for something new to grow. We didn’t need a whole harvest. We just needed the right thing at the right time. It wasn’t perfect, and it wasn’t permanent, but it got things back on track, and gave us space to understand and solve the bigger problem. Not only did we create conditions for the immediate fix, but we were able to leave the long term crop growing in its healthy soil.
So the next time you’re looking at your incident management process, I wonder if you will ask yourself this: Does this create the conditions for creative, effective, timely incident response? Does this renew my soil, or does it deplete it?
메타데이터
- post_id
- fc00b1f9c7eb
- slug
- incident-management-farming-in-a-crisis-fc00b1f9c7eb
- url
- https://medium.com/@Spamaps/incident-management-farming-in-a-crisis-fc00b1f9c7eb
- canonical_url
- https://medium.com/@Spamaps/incident-management-farming-in-a-crisis-fc00b1f9c7eb
- author_url
- https://medium.com/@Spamaps
- status
- ok
- fetched_at
- 2026-06-23 03:48:11