How to Run I&O Like a Product Organization: A Practitioner’s Guide to the Five Capabilities That…
There’s a conversation happening in I&O leadership circles that doesn’t show up in most vendor briefings.
How to Run I&O Like a Product Organization: A Practitioner’s Guide to the Five Capabilities That Now Define Infrastructure Leadership

There’s a conversation happening in I&O leadership circles that doesn’t show up in most vendor briefings.
It’s not about which cloud provider to standardize on, or whether to adopt a new observability platform. It’s about a more fundamental shift in how infrastructure and operations teams are being evaluated — and what that means for every decision they make.
I&O is being held to a product standard.
Not in the abstract sense of “treat your users like customers.” In the concrete sense of being judged on the same dimensions product teams are judged on: customer experience, speed to recovery, reliability and resilience, cost discipline, continuous improvement, and measurable outcomes. All at the same time. On a platform that is becoming more fragmented, more dynamic, and more expensive every year.
If you lead an infrastructure or operations team and that framing doesn’t produce some productive discomfort, it probably means you haven’t felt the full weight of it yet.
This piece builds on five signals I observed across analyst sessions, sponsor conversations, and peer discussions at IOCS 2025. The full synthesis — including the common thread connecting all five — is in my LinkedIn article here: https://www.linkedin.com/pulse/five-signals-from-iocs-2025-io-leaders-cant-ignore-salil-kulkarni-lnahe. What I want to do here is go deeper on what the “I&O as product organization” shift actually requires in practice: the capabilities, the accountability model, and where most organizations are still building on a foundation that won’t hold.
What “I&O as Product Organization” Actually Means
When product teams are evaluated, they don’t get credit for effort. They get credit for outcomes. A product that ships on time but breaks in production isn’t a success. A product with beautiful architecture that no one uses isn’t a success.
The same logic is now being applied to I&O.
“We kept the lights on” is no longer the bar. The bar is: How fast did you recover? What was the business impact of the outage? How did you change the blast radius of that infrastructure update? What did that cost, and was it the right cost?
That’s a different accountability model than most I&O organizations were built for. And it requires a different set of platform capabilities than most I&O organizations currently have.
Five capabilities surfaced consistently at IOCS 2025 as the ones that define whether an I&O team can actually operate at product-organization standards. Here’s what each requires — and where the real gaps are.
Capability 1: AI That Operates Under Governance, Not Just Under Instruction
The tone around AI in I&O has shifted. A year ago, the conversation was “what can AI do for IT operations?” Today the conversation is “who is accountable when an AI agent makes a mistake, and what stops it from causing a cascading failure?”
That’s a maturity shift worth taking seriously.
AI is moving from pilot to policy. And the organizations that are moving fastest aren’t the ones with the most AI features deployed. They’re the ones that have built a governance model around AI actions: defined accountability for agent decisions, guardrails that prevent automation from compounding a failure, audit trails that let teams understand what the AI did and why.
What it requires in practice: AI actions in IT operations need to be bounded, traceable, and reversible where possible. That means change workflows with approval gates for high-risk AI-initiated actions. It means event correlation that surfaces AI reasoning, not just AI conclusions. It means human-in-the-loop design for actions above a defined impact threshold — not because AI can’t handle them, but because governance requires it.
Where most organizations fall short: Treating AI governance as a compliance exercise rather than an operational design problem. Governance that lives in policy documents but not in the actual workflow — where the AI can still take consequential actions without an auditable decision trail.
The product parallel: a product team that ships features without instrumentation doesn’t know when things break. An I&O team that deploys AI without governance doesn’t know when it breaks things.
Capability 2: Infrastructure Visibility That Precedes Every Optimization Decision
Infrastructure economics are breaking. That was one of the clearest signals at IOCS 2025 — and it’s not primarily a technical problem.
AI workloads demand data locality and predictable latency, which is forcing honest conversations about cloud placement decisions made three years ago under different assumptions. Licensing changes are forcing recalculation of costs that were once predictable. Cloud bills are harder to forecast as consumption models evolve.
The uncomfortable truth underneath all of this: most organizations are trying to optimize environments they don’t fully understand. They’re accumulating infrastructure — adding cloud regions, adding tools, adding hybrid execution layers — without a corresponding improvement in visibility into what exists, how it connects, and what it costs to run.
What it requires in practice: Infrastructure strategy and service operations have to converge. “What belongs where and why?” is not a question you can answer from a spreadsheet or a quarterly asset report. It requires continuously updated visibility into assets, dependencies, and the cost and performance implications of where workloads live. IT Discovery — real continuous discovery, not periodic scanning — becomes a strategic input to infrastructure decisions, not an implementation detail.
Where most organizations fall short: Treating discovery as a CMDB population exercise rather than a decision-support capability. Running infrastructure optimization workshops with data that is months old. Making cloud placement decisions based on architecture diagrams that don’t reflect the environment as it currently exists.
The product parallel: you can’t build a product roadmap without understanding how your current product is being used. You can’t build an infrastructure strategy without understanding how your current infrastructure is being used.
Capability 3: Resilience Built on Service Context, Not Security Perimeter
Cyber resilience sounded different at IOCS 2025. Less like a security conversation, more like an operations conversation.
That’s because the question has changed. In a distributed enterprise, resilience isn’t primarily about blocking threats — it’s about maintaining service continuity when things go wrong and recovering fast when they don’t. And fast recovery requires knowing what depends on what, where failure propagates, what to restore first, and what changes introduced risk.
None of those questions can be answered by a security tool. They can only be answered by someone with an accurate, current understanding of service topology.
What it requires in practice: Resilience planning needs service context as its foundation. That means dependency-aware recovery runbooks — where restoration sequences are informed by actual current topology, not assumed topology from the last documentation refresh. It means change risk analysis that draws on live relationship data, so that the blast radius of a proposed change is calculated against the environment as it exists today. It means incident response workflows where the first question — “what depends on this, and what’s at risk?” — can be answered in seconds, not minutes.
Where most organizations fall short: Resilience plans that were accurate when written but haven’t been validated against a changed environment. Recovery runbooks that assume service dependencies that no longer exist or miss dependencies that have been added. Cyber recovery exercises that test the plan but not the topology data the plan depends on.
The product parallel: a product incident response plan that assumes an old architecture will fail the first time it encounters the real one.
Capability 4: Cost and Operating Model Decisions Made Together, Not Sequentially
Here’s the shift that was most striking at IOCS 2025: cost conversations are happening inside every I&O conversation, not as a separate downstream discussion.
That’s new. For most of the last decade, the pattern was: I&O makes technical decisions, finance reviews the cost implications afterward. The new pattern is: cost, sustainability, and operating model are constraints that shape the technical decision in real time.
AI introduces new consumption models that are harder to forecast than traditional infrastructure. Observability data volumes are growing faster than budgets. Tool sprawl is creating redundant spend that’s increasingly hard to justify. Talent constraints are accelerating the push toward automation — but automation built on fragile foundations creates its own operational cost.
What it requires in practice: Every tool decision needs to be evaluated as an operating model decision and a cost model decision simultaneously. That means understanding not just the licensing cost of a tool, but the operational overhead it creates, the data volume it generates, and how it integrates with the rest of the environment. It means sustainability showing up as a real constraint — energy consumption, data center efficiency, workload placement — not as a separate ESG reporting exercise.
Where most organizations fall short: Tool procurement processes that evaluate capabilities without modeling operating costs. Observability deployments that grow data volumes without corresponding growth in signal quality. Automation investments that reduce labor cost in one area while creating fragility cost in another.
The product parallel: product teams that optimize for feature velocity without modeling technical debt are building a cost problem they’ll pay for later. I&O teams that optimize for capability without modeling operational cost are doing the same thing.
Capability 5: Observability That Produces Decisions, Not Dashboards
The observability conversation at IOCS 2025 was not about more visibility. It was about less noise.
Organizations aren’t asking for more dashboards. They’re asking for fewer tools, faster diagnosis, correlated insights, and clear service-level impact. They want observability to function as a decision engine — not a data repository with a visualization layer on top.
The challenge is that AI-driven observability without service context becomes expensive noise. More telemetry correlated by more models still produces confusion if the underlying question — “which service does this belong to, and what depends on it?” — can’t be answered accurately.
What it requires in practice: Observability has to be grounded in accurate service mapping and relationship data. Telemetry needs to be tied to services, dependencies, and business impact — not just to infrastructure components. That means the observability platform and the CMDB/service map need to be genuinely integrated, not just loosely connected through a one-time import. When an anomaly fires, the system should be able to answer: what service is affected, what are its upstream and downstream dependencies, what is the business impact, and what changed recently that might have caused this.
Where most organizations fall short: Observability deployments that are rich in telemetry but poor in context. Platforms that can tell you a node is degraded but can’t tell you which business service is at risk. AI correlation that produces probable root causes pointing to components whose relationships to current services haven’t been validated since the last infrastructure change.
The product parallel: analytics that tell you features are being used but can’t tell you whether users are achieving their goals produce activity metrics, not outcome metrics. Observability that tells you infrastructure is degraded but can’t tell you which services are affected produces event metrics, not outcome metrics.
The Capability That Underlies All Five
Here’s what connects the five signals: every one of them depends on a continuously updated understanding of what exists in the environment and how it connects.
AI governance requires knowing what the AI acted on and what the impact was — which requires accurate service context. Infrastructure optimization requires knowing what exists before you can decide what belongs where. Resilience requires knowing what depends on what before you can recover in the right order. Cost modeling requires knowing what your environment actually looks like before you can model what it costs to run. Observability requires service context to produce signal rather than noise.
The CMDB was supposed to be that foundation. Most CMDBs weren’t designed for environments that change as fast as modern hybrid, multicloud architectures do. They reflect the environment as it was at the last scan, not the environment as it is now.
The organizations at IOCS 2025 that were furthest along on all five capabilities had one thing in common: they had invested in continuous, automated discovery and dynamic service mapping as a foundation — not as a feature to add on top of an existing toolset.
If I&O is a product organization, service visibility is the platform the product runs on.
The full synthesis from IOCS 2025 — including the five signals and the deeper challenge underneath them — is in my LinkedIn article here: https://www.linkedin.com/pulse/five-signals-from-iocs-2025-io-leaders-cant-ignore-salil-kulkarni-lnahe. If you’re navigating any of these capability questions, I’d welcome the chance to compare notes.
메타데이터
- post_id
- 2e6ab966bfd8
- slug
- how-to-run-i-o-like-a-product-organization-a-practitioners-guide-to-the-five-capabilities-that-2e6ab966bfd8
- url
- https://medium.com/@saliljk/how-to-run-i-o-like-a-product-organization-a-practitioners-guide-to-the-five-capabilities-that-2e6ab966bfd8
- canonical_url
- https://medium.com/@saliljk/how-to-run-i-o-like-a-product-organization-a-practitioners-guide-to-the-five-capabilities-that-2e6ab966bfd8
- author_url
- https://medium.com/@saliljk
- status
- ok
- fetched_at
- 2026-06-15 20:49:13