← Back to list

AI, Climate, and Infrastructure Risk: The New Governance Test

The next frontier in risk leadership is connecting digital ambition to the power, physical systems, markets, and decision rights that make…

Kevin Pausicles · 2026-06-18 12:19 · 0 claps · 47.3 min read
#ai-risk #climate-risk #infrastructure-resilience #energy-market #governance-risk
Open on Medium ↗
Wiki topics: BIZ · Business Strategy ECO · Economy · General 🚀 · Self Improvement

AI, Climate, and Infrastructure Risk: The New Governance Test

The next frontier in risk leadership is connecting digital ambition to the power, physical systems, markets, and decision rights that make it possible.

In one executive meeting, the organization approves an ambitious artificial intelligence strategy.

The presentation covers productivity, customer experience, automation, data quality, model performance, cybersecurity, privacy, and responsible use. The business case is compelling. The governance structure appears mature. The roadmap is aggressive but credible.

In another meeting, operations reviews business continuity.

In another, sustainability presents climate scenarios.

Treasury discusses energy costs and market exposure. Procurement reviews cloud contracts. Technology evaluates model providers. Facilities assess backup power. Human resources consider workforce readiness. Legal updates policy. The board receives separate reports from separate teams.

Every one of those discussions can be competent.

Taken together, they can still miss the risk.

The problem is not that organizations lack specialists. The problem is that the most important exposures increasingly sit between their areas of expertise.

AI is usually categorized as a technology, model, cyber, data, or conduct risk.

Climate is often categorized as a sustainability, disclosure, regulatory, or long-term strategic risk.

Power and physical infrastructure are frequently treated as operating assumptions.

Energy-market volatility is assigned to finance, procurement, trading, or treasury.

Operational resilience is managed through continuity plans, incident procedures, and recovery targets.

Those categories remain useful. But the boundaries between them are becoming less useful.

AI growth creates new claims on computing capacity, electricity, cooling, telecommunications, specialist hardware, cloud infrastructure, and highly concentrated vendors. Climate conditions affect the availability, performance, location, and cost of those same physical systems. Energy markets transmit stress between supply, demand, weather, infrastructure, and price. Operational processes become dependent on digital services whose resilience may be determined far outside the organization’s direct control.

The result is a connected exposure chain:

AI strategy depends on digital infrastructure. Digital infrastructure depends on physical infrastructure. Physical infrastructure depends on energy, water, communications, supply chains, people, and place. Those dependencies are exposed to weather, markets, regulation, concentration, and operational failure.

The International Energy Agency reported that electricity demand from data centers grew by 17 percent in 2025 and projects that total data-center electricity consumption could double by 2030, while consumption from AI-focused facilities could triple. At the same time, it has identified grid connections, transformers, gas turbines, chips, and other equipment as emerging physical bottlenecks. These forecasts are uncertain, as any serious forecast must be, but their direction is already relevant to governance.

In the United States, the Department of Energy cites research estimating that data centers consumed about 4.4 percent of national electricity in 2023 and could account for approximately 6.7 to 12 percent by 2028. The national percentage tells only part of the story. The operational consequences are local and regional: a large new load does not connect to an abstract national grid. It connects at a particular location, through particular transmission and distribution assets, within a particular resource-adequacy and market structure.

This does not mean every organization adopting AI needs to become a power-system operator.

It does mean every organization making AI strategically important must understand that digital ambition is a claim on physical capacity.

The claim may be direct, as it is for data-center developers, utilities, energy companies, industrial operators, and infrastructure providers. Or it may be transferred through cloud providers, software vendors, managed-service companies, telecommunications networks, and outsourcing arrangements.

Transfer is not elimination.

The cloud has always been physical. AI is making that fact harder to ignore.

Climate risk creates a similar governance challenge. A climate scenario is not useful because it produces a sophisticated chart. It is useful when it tells management which assets, services, suppliers, communities, employees, contracts, or financial assumptions become vulnerable — and what decisions must then be made.

Heat is an operating condition.

Flood is an operating condition.

Drought, wildfire, severe wind, extreme cold, storm surge, water stress, and changing seasonal patterns are operating conditions.

So are insurance availability, infrastructure maintenance, asset life, workforce safety, logistical access, energy demand, and supply interruption.

This is not a political framing. It is a risk framing.

The organization does not have to settle every public debate before deciding whether a critical facility can withstand a flood, whether a supplier has a viable alternate route, whether backup generation will start, whether a cloud region shares an infrastructure dependency with its supposed alternative, or whether decision rights are clear when capacity becomes constrained.

That is the new governance test.

It is not whether the organization has an AI policy.

It is not whether it publishes climate disclosures.

It is not whether operations has a continuity plan.

It is not whether treasury monitors prices.

It is whether leadership can connect ambition to operating reality before a disruption connects it by force.

1. The risk map has changed

Most corporate risk frameworks are organized by nouns.

Cyber risk. Model risk. Climate risk. Market risk. Operational risk. Third-party risk. Technology risk. Conduct risk. Liquidity risk. Strategic risk.

Demand accelerates.

A grid connection is delayed.

A transformer fails.

A heatwave increases load while reducing system flexibility.

A storm damages transmission, distribution, communications, transport, or supplier access.

A vendor allocates scarce capacity.

An AI service degrades rather than disappearing completely.

A critical workflow has no tested manual alternative.

Energy prices gap.

Collateral requirements rise.

Insurance terms change.

Employees cannot safely reach a site.

Customers need more support at exactly the moment automated support becomes less reliable.

The event moves through the organization faster than its risk taxonomy does.

This is why the risk map must evolve. The categories themselves are not wrong; the interactions are under-governed.

For years, many organizations could treat digital infrastructure as an abundant utility. Computing capacity could be purchased when required. Cloud expansion appeared close to frictionless. Power availability was a problem for the vendor or the utility. Weather was incorporated into facility planning, emergency management, and insurance. Energy prices were a budget or hedging issue. AI use was narrow enough to remain inside specialist analytical processes.

That operating environment is changing.

AI is moving from isolated use cases into the fabric of work. It is being introduced into software development, customer support, fraud detection, maintenance, forecasting, scheduling, document processing, research, procurement, compliance, engineering, security, trading support, knowledge management, and executive decision processes.

The risk profile changes when a tool becomes a workflow.

It changes again when a workflow becomes a control.

It changes again when the control becomes difficult to operate without the tool.

An experimental assistant can be unavailable without serious consequence. An AI-enabled customer channel, operational decision engine, diagnostic process, scheduling capability, or control-monitoring function may create a very different tolerance for failure.

The key question is not simply, “Do we use AI?”

It is:

What has the organization become unable to do safely, lawfully, economically, or at sufficient scale without it?

That is an operational-resilience question, not only a model-governance question.

Climate risk is following the same path. It is moving from an external scenario into the performance of assets and services. Organizations may still discuss it through long-range pathways, reporting obligations, carbon assumptions, or policy developments. Those matters can be important. But the immediate governance value comes from translating climate information into operating consequences.

Which facilities face increasing heat stress?

Which transmission, transport, telecommunications, water, or supplier dependencies are geographically concentrated?

Which sites are protected against a historical hazard level that may no longer represent the relevant operating range?

Which processes rely on employees working outdoors or inside poorly cooled facilities?

Which products become more expensive to insure?

Which suppliers have recovery plans that assume the same regional infrastructure on which the organization itself depends?

Which assets have a nominal life extending into conditions materially different from those used in their original design?

Which adaptation investments have been deferred because no single budget owner captures the full benefit?

These are not abstract climate questions. They are questions about service continuity, cost, capital, liability, and leadership.

The energy system sits between AI ambition and climate exposure. It translates physical scarcity into operational constraints and market signals. Electricity demand, generation availability, fuel supply, transmission capacity, interconnection, weather, maintenance, storage, and regulatory requirements interact continuously. The system does not care which corporate department owns which risk category.

Neither does a disruption.

The risk leader’s job is therefore changing.

The next serious frontier is not better administration of separate risk domains. It is the ability to identify, govern, and rehearse connected risk across digital systems, physical infrastructure, environmental conditions, markets, third parties, and executive decisions.

That does not require the chief risk officer to take over technology, sustainability, operations, finance, or procurement. It requires the risk function to create a view that none of those functions can create alone.

The organization needs to know:

  • What are our critical services?
  • What physical and digital dependencies support them?
  • Where are those dependencies concentrated?
  • Under what conditions do they degrade?
  • What breaks first?
  • What decisions become necessary?
  • Who has the authority to make them?
  • What service level are we prepared to protect?
  • What activity would we stop first?
  • What residual risk has leadership knowingly accepted?

Without that connected view, an organization can have strong controls in every individual domain and remain fragile at the points where the domains meet.

2. Why AI risk is becoming physical

AI governance has developed rapidly around legitimate concerns: model performance, bias, explainability, security, privacy, intellectual property, data quality, hallucination, human oversight, legal compliance, and responsible use.

Those controls are necessary.

They are not sufficient.

AI does not operate in a policy document. It operates through hardware, data centers, electricity, cooling, networks, software libraries, cloud regions, APIs, identity systems, data pipelines, specialist staff, and third-party contracts.

A model may be digital. Its operating system is physical and institutional.

Compute is not an abstract resource

The phrase “compute demand” can sound technical and remote. In practice, it means servers, processors, memory, storage, network capacity, cooling systems, buildings, electrical interconnections, generation, transmission, distribution, backup systems, maintenance, and supply chains.

It also means location.

Two workloads with the same total energy consumption can have different risk profiles because they run at different times, require different levels of availability, use different facilities, depend on different network paths, or concentrate demand in different regions.

The distinction between training and inference matters. So does the distinction between batch processing and real-time decision support. Some workloads can be delayed or shifted. Others are expected to respond immediately. Some can tolerate reduced performance. Others support a process in which latency, interruption, or degraded output creates customer, financial, safety, or compliance consequences.

This means that “How much power does AI use?” is only a starting question.

A risk leader should also ask:

What is the load profile?

How flexible is it?

Where is it served?

What happens during capacity constraints?

Who gets priority?

What contractual rights govern allocation?

What service degradation should we expect before a complete outage?

Can workloads move to another region without creating data, latency, sovereignty, cost, or control problems?

Does the alternate region depend on the same transmission corridor, cloud control plane, network provider, identity service, or model company?

The physical risk does not disappear because the infrastructure is outsourced. It becomes a third-party dependency, often with reduced transparency.

Logical diversity can conceal physical concentration

An organization may believe it has diversified because it uses multiple applications.

Yet those applications may rely on the same foundation model, cloud provider, identity platform, chip architecture, content-delivery network, data source, or regional infrastructure.

Three vendors can be three separate names on a procurement list while representing one underlying dependency.

The same problem appears inside AI supply chains. A company may build its own interface but rely on an external model. It may use several models hosted on one cloud. It may contract with separate providers that rely on common open-source libraries or shared infrastructure. It may retain access to an alternate model that has never been tested at the required volume.

This is common-mode risk: apparent redundancy that fails under the same stress.

Concentration risk is already receiving attention from public authorities in sectors where operational dependencies have systemic consequences. A U.S. Treasury review of AI in financial services, for example, included monitoring concentration risk among its policy considerations. The broader lesson applies well beyond finance: an organization needs visibility into the common infrastructure beneath its nominally separate AI services.

Vendor reviews often focus on financial condition, cybersecurity, privacy, contract terms, and service levels. Connected AI governance must go further.

Where does the service run?

Which components are subcontracted?

What are the provider’s own critical dependencies?

Can the provider change models, infrastructure, or terms without meaningful customer control?

What happens to output quality during failover?

How is capacity allocated during widespread demand?

Can the organization retrieve its data, prompts, configurations, logs, and evaluation records quickly enough to migrate?

Does the alternate provider actually support the required use case, jurisdiction, volume, and control environment?

A contract can assign responsibility. It cannot manufacture unavailable capacity during a crisis.

Workflow dependency can grow faster than governance

The first stage of AI adoption is usually experimentation.

Employees use a tool to summarize documents, improve drafts, query information, or assist with analysis. Failure is inconvenient but manageable.

The second stage is integration.

AI becomes embedded in business applications, development environments, customer channels, monitoring systems, planning tools, or internal knowledge platforms.

The third stage is dependency.

Staffing models, turnaround times, service promises, control processes, and customer expectations begin to assume the tool is available.

The fourth stage is institutional atrophy.

Manual capacity is reduced. Skills weaken. Legacy processes are retired. New employees are trained on the AI-enabled workflow rather than the process beneath it. Data is structured for machine use but becomes harder for people to interpret quickly. Performance targets are reset around automated throughput.

At that point, the organization may still describe the tool as an assistant while operating as though it were infrastructure.

This is where fallback risk becomes serious.

A manual process documented eighteen months ago may no longer be viable. The people named in the procedure may have moved roles. The relevant access permissions may have expired. The volume may have doubled. Source data may no longer be presented in human-readable form. The organization may have enough manual capacity for an isolated outage but not for a regional disruption combined with higher customer demand.

Fallback should therefore be measured, not asserted.

How much of the critical workload can be completed without the AI service?

For how long?

At what error rate?

With which staff?

Under what demand assumptions?

What is the minimum safe service?

Which customers or processes receive priority?

What information must be preserved so that people can take over?

How long does it take to activate the fallback?

Has it been tested during normal operations rather than only discussed in a workshop?

The most dangerous failure may not be a clean outage. It may be partial degradation.

A system that stops completely creates a visible incident. A system that becomes slower, less accurate, less consistent, or subtly different after a model update can remain inside the workflow while degrading decisions.

That creates difficult governance questions.

When does reduced model quality become an operational incident?

Who has the authority to withdraw the system?

How quickly can the organization detect a distribution shift, capacity constraint, vendor change, or control failure?

Can it compare current performance with an approved baseline?

Does it understand how a fallback model changes outcomes?

Is “available” being confused with “fit for purpose”?

AI availability and AI integrity are connected

Traditional resilience planning often separates availability from quality. A system is either running or unavailable.

AI complicates that distinction.

A service may be technically available but produce materially different outputs because of model changes, altered retrieval data, reduced context, capacity management, new safety filters, latency, integration errors, or upstream data problems.

A failover arrangement may restore access without restoring equivalent behavior.

That matters when the model supports customer treatment, operational decisions, coding, forecasting, maintenance, security, or controls.

The organization therefore needs at least three operating states:

Normal service: The approved model, data, controls, capacity, and monitoring are functioning within tolerance.

Degraded service: The AI capability remains available, but accuracy, latency, scope, data access, or control assurance has moved outside normal tolerance.

Fallback service: The organization has deliberately shifted to an alternate model, manual process, reduced service, or controlled suspension.

Those states require different decisions. They also require thresholds.

A generic policy stating that humans remain accountable does not answer how a human is supposed to intervene when volume exceeds manual capacity, the degradation is difficult to observe, or the relevant expertise has been allowed to decline.

Human oversight must be designed into the operating model.

Model governance must connect to infrastructure governance

The NIST AI Risk Management Framework organizes AI risk activity around four functions — govern, map, measure, and manage — and emphasizes continuous risk management across the AI lifecycle. That lifecycle perspective is valuable because it discourages organizations from treating approval as the end of governance.

Connected risk governance extends the lifecycle view into the physical and operational environment.

To govern an AI system, the organization must know not only what the model does, but what business service depends on it.

To map risk, it must identify not only model limitations, but infrastructure, data, vendor, geographic, workforce, and energy dependencies.

To measure risk, it must track not only accuracy and bias, but availability, latency, capacity, concentration, change, fallback readiness, and recovery.

To manage risk, it must establish decision rights for degraded operation, vendor substitution, workload reduction, manual intervention, and temporary suspension.

AI governance that stops at the model boundary is incomplete.

The relevant unit of risk is the AI-enabled service, including everything required to deliver it safely.

3. Why climate risk is becoming operational

Climate risk is sometimes weakened by the way organizations frame it.

When it is treated only as a values issue, people debate intent.

When it is treated only as a disclosure issue, people debate wording.

When it is treated only as a distant scenario, people debate assumptions.

An operating organization needs a more direct question:

What can the physical environment do to our ability to deliver?

That question is concrete enough to govern.

Hazard is not the same as risk

A disciplined climate-risk assessment separates four elements:

Hazard: The physical event or condition, such as heat, flood, wildfire, drought, severe wind, extreme cold, or sea-level change.

Exposure: The assets, people, suppliers, infrastructure, services, or communities located where the hazard can affect them.

Vulnerability: The extent to which those exposed elements can withstand, adapt to, or recover from the hazard.

Consequence: The operational, financial, legal, safety, customer, market, or strategic result.

This distinction matters because a hazard map alone does not tell management what to do.

Two facilities can face similar heat and have different risk because one has adequate cooling, backup power, maintenance, staffing, and demand flexibility while the other does not.

Two suppliers can face the same flood zone and have different recovery capability because one has alternate production and transport routes while the other relies on a single site.

Two data centers can face the same regional weather event but differ in grid connection, water dependency, equipment design, contractual priority, or access to backup fuel.

The purpose of climate-risk management is not merely to identify hazards. It is to reduce vulnerability and consequence.

Climate impacts move through infrastructure systems

The IPCC has emphasized that infrastructure is susceptible to compounding and cascading risks. A single event can propagate across interconnected systems and amplify impact across locations and time.

That description is directly relevant to corporate resilience.

Consider extreme heat.

Heat can increase electricity demand for cooling. It can affect workforce safety and productivity. It can constrain some equipment. It can raise water requirements. It can increase wildfire conditions. It can place additional stress on transport and communications infrastructure. It can occur across a broad region, reducing the value of nearby alternatives that share the same conditions.

Now combine heat with an AI-enabled operating model.

Customer demand may rise. Automated systems may process higher volumes. Data-center cooling and power demand may increase. Grid conditions may tighten. Electricity-market prices may move. A cloud or infrastructure provider may initiate capacity-management procedures. Employees may be unable to work safely at full productivity. The manual fallback may therefore be least available when it is most needed.

The relevant risk is not “heat” in one register and “AI outage” in another.

It is the interaction.

The same applies to flood. Flood can affect facilities, substations, telecommunications, transport, employee access, fuel delivery, supplier operations, and community services at the same time.

Wildfire can damage infrastructure directly, create smoke and air-quality hazards, interrupt transmission, close roads, prompt preventive shutdowns, and displace staff.

Drought can affect water-dependent operations, hydropower availability, agriculture, transport, and thermal-generation cooling, while also altering commodity and energy conditions.

Extreme cold can raise demand, affect fuel delivery, damage equipment, and test winterization assumptions across a wide geography.

Storms can create simultaneous demand for emergency response, customer communication, field repair, logistics, temporary accommodation, claims handling, and liquidity.

The organizational question is not whether each hazard appears in a risk register. It is whether management understands the sequence through which the hazard becomes a service failure.

Disclosure is not resilience

Climate disclosure can improve transparency and force useful analysis. It can also create false comfort if the organization confuses reporting with preparedness.

A scenario in a report may describe higher temperatures, changing precipitation, or increased hazard intensity. An operating model must translate that into questions such as:

Which critical service is affected?

Which site, asset, supplier, or workforce dependency is exposed?

At what threshold does performance decline?

What monitoring provides warning?

What adaptation has been completed?

Which control is temporary?

What investment is required?

Who owns the decision?

What happens if the planned adaptation is delayed?

What residual risk is being accepted?

A disclosure can state that an organization faces flooding risk. Resilience requires knowing whether pumps are maintained, barriers are adequate, generators sit above the relevant water level, fuel is accessible, communications are redundant, staff can reach the site, and recovery priorities are agreed.

A disclosure can state that heat may affect operations. Resilience requires workforce thresholds, cooling capability, equipment limits, load assumptions, demand-management plans, supplier contingencies, and decision rights.

A disclosure can state that insurance costs may rise. Financial resilience requires understanding deductibles, exclusions, limits, insurer concentration, uninsured loss, repair inflation, business-interruption assumptions, and the capital consequences of a prolonged event.

The difference is execution.

Climate scenarios must reach the operating horizon

Some organizations separate time horizons too sharply.

Operational risk focuses on the next quarter or year.

Strategic risk looks three to five years ahead.

Climate scenarios may extend decades.

This can create a governance gap. The near-term team assumes the risk is too distant. The long-term team produces analysis too broad for today’s decisions.

Yet many decisions made now have long operating lives.

A facility lease, data-center contract, energy arrangement, infrastructure investment, insurance strategy, supply agreement, equipment purchase, network design, or site selection can create exposure for years.

The correct question is not, “Will this climate scenario occur before our next planning cycle?”

It is:

Does the decision we are making today create a dependency that will remain in place as conditions change?

A twenty-year scenario can be relevant to a five-year contract if migration takes four years.

A ten-year hazard trend can be relevant to equipment bought today.

A chronic increase in hot days can matter immediately if the current system already operates close to thermal limits.

A changing insurance market can affect project economics well before physical damage occurs.

Time horizon should therefore be linked to decision life, not only forecast year.

Physical risk and transition risk affect one operating model

Organizations often separate physical climate risk from transition risk.

Physical risk concerns weather and longer-term environmental conditions.

Transition risk concerns changes in policy, technology, customer behavior, legal expectations, financing, and market structure.

The distinction is analytically useful, but management must consider how both affect the same decisions.

An asset may face physical exposure while also requiring capital to meet new operating or regulatory standards.

A supplier may face higher energy and insurance costs while being asked to change technology.

A company may need to increase resilience spending at the same time as it funds AI infrastructure and other strategic priorities.

A power system may face rapid demand growth while generation, transmission, permitting, fuel, and market structures are also changing.

The governance challenge is not choosing which risk matters. It is managing the portfolio of constraints.

This is where risk appetite becomes practical.

How much service interruption will the organization tolerate?

How much concentration will it accept?

How much capital will it commit to adaptation?

Which locations or suppliers require exit plans?

Which strategic initiatives should be sequenced rather than launched simultaneously?

What assumptions must be true for the plan to remain credible?

When should management slow expansion because infrastructure readiness is lagging?

These decisions are not made effectively when climate risk sits apart from strategy and operations.

Climate risk becomes useful when it changes a decision.

4. Energy markets sit at the intersection

Energy and power markets provide one of the clearest views of connected risk because they make physical conditions economically visible.

AI requires electricity.

Electricity systems must balance supply and demand continuously.

Weather affects both sides of that balance.

Infrastructure determines what can move, where it can move, and when.

Markets translate scarcity, congestion, fuel availability, resource performance, and demand into prices and operating signals.

A company does not need to trade power to be exposed to this system.

It may face exposure through utility rates, contracts, cloud pricing, supplier costs, project delays, service availability, capacity charges, backup requirements, customer demand, or regional economic conditions.

Electricity availability is a strategic input

For many businesses, electricity has historically appeared as a reliable line item: important, but not central to strategic design.

AI changes that assumption for some organizations directly and for many others indirectly.

The strategic question is no longer only whether power can be purchased. It is whether sufficient capacity is available at the required location, on the required timeline, with the required reliability, at a cost consistent with the business model.

A technically feasible AI plan can therefore be operationally infeasible.

The model may work.

The use case may be approved.

The data may be ready.

The vendor may be selected.

The organization may still face a constrained grid connection, delayed equipment, limited regional capacity, rising cost, or a mismatch between the speed of digital deployment and the speed of infrastructure delivery.

The IEA’s recent analysis describes exactly this tension: rapid data-center growth is encountering physical bottlenecks in grid connections and key equipment even as computing efficiency continues to improve.

Efficiency matters, but it does not automatically reduce total demand. Lower energy per task can be offset by more users, larger models, autonomous agents, new applications, and higher volumes.

For governance purposes, this creates two forms of uncertainty:

Demand uncertainty: How quickly will AI use grow, and how much computing will each business process require?

Supply uncertainty: How quickly can electricity, networks, equipment, and facilities be made available?

The organization can be wrong in either direction.

It can overbuild for demand that does not materialize.

It can underprepare for adoption that moves faster than infrastructure.

Risk leadership should not pretend to remove that uncertainty. It should ensure the organization has options.

Market prices expose physical assumptions

Energy prices are not a complete measure of resilience, but they are an important signal.

FERC reported that wholesale electricity prices rose in most U.S. regions in 2025, increasing 26 percent year over year nationwide, while electricity demand also grew substantially. It identified electrification and data-center expansion among the factors prompting expectations of higher future load and new approaches to resource adequacy and large-load integration.

A single annual average should not be overinterpreted. Power-market exposure is shaped by location, timing, contract structure, fuel mix, transmission, weather, and the difference between average and peak conditions.

That is precisely the point.

An AI strategy based only on annual electricity consumption misses operational shape.

When does the workload run?

Can it be shifted?

Is it concentrated during peak periods?

Does it require uninterrupted service?

What happens during scarcity pricing?

Is cost passed through by a vendor?

Does the organization have a fixed-price contract that protects budget but not availability?

Does an alternate location face different price risk but the same equipment constraint?

How do energy costs affect the economics of marginal AI use cases?

How does volatility affect suppliers whose financial resilience may be weaker than the organization’s own?

Market risk is not limited to price direction. Relevant exposures can include:

  • Volume risk
  • Timing and shape risk
  • Regional basis and congestion risk
  • Fuel-price exposure
  • Capacity and resource-adequacy risk
  • Counterparty risk
  • Credit and collateral requirements
  • Liquidity risk during stressed markets
  • Contract renewal and pass-through risk
  • Regulatory and permitting risk
  • Asset-performance risk
  • Supplier financial risk

These exposures interact.

A period of severe weather can increase demand and prices while disrupting physical operations. A supplier may face higher energy cost and lower production simultaneously. A service provider may maintain availability but pass through higher cost at renewal. A grid constraint may make one region expensive while another remains relatively stable, exposing assumptions hidden by national averages.

The organization needs to understand which variables affect service, which affect cost, and which affect both.

Reliability is local, temporal, and conditional

It is easy to speak about generation capacity in aggregate. Operating reliability depends on more than aggregate nameplate capacity.

Resources have different operating characteristics.

Transmission constraints matter.

Fuel availability matters.

Maintenance and forced outages matter.

Weather affects demand and resource performance.

The timing of supply must match the timing of demand.

Interregional transfer capability matters.

Forecast error matters.

So do demand response and operating reserves.

NERC’s current reliability work reflects the challenge of rapidly growing loads. Its 2026 summer assessment, for example, notes that rising data-center forecasts are contributing to load growth in the Midcontinent region and could increase future reliability risk if resource additions fail to keep pace.

This should not be read as a prediction of inevitable failure. Reliability assessments are scenario-based and depend on assumptions. It should be read as a governance signal: digital demand and physical resource planning are moving on different timelines, and the gap requires active management.

For a board or executive team, the relevant questions are not technical trivia.

Where is our strategy dependent on new electrical capacity?

What is the credible connection timeline?

What assumptions have been made about resource availability?

Which delays would affect the business case?

What obligations remain if capacity arrives late?

What flexibility has been designed into location, scheduling, contracts, or workload?

Who monitors changes in the underlying infrastructure plan?

At what point does management revise the digital roadmap?

Energy matching and operational resilience are different questions

Organizations may seek to match electricity use with contracted generation or environmental attributes. That can support legitimate environmental and commercial objectives.

It does not answer every resilience question.

Annual energy matching does not necessarily establish that electricity is available at the required hour and location.

A long-term contract does not necessarily eliminate regional congestion.

A generation commitment does not necessarily provide the same service during extreme conditions.

Backup generation does not guarantee resilience if fuel, maintenance, testing, cooling, emissions limits, or start reliability are weak.

Battery storage can provide valuable flexibility but has duration and operating constraints.

On-site generation can reduce some dependencies while introducing equipment, fuel, maintenance, permitting, safety, and concentration risks.

The risk function does not need to advocate one technology. It needs to make sure the organization distinguishes objectives.

Is the objective cost stability?

Availability?

Carbon performance?

Speed to connection?

Regulatory compliance?

Operational independence?

Market participation?

Community acceptance?

No single arrangement necessarily optimizes all of them.

Good governance makes the trade-offs explicit.

AI can strengthen the energy system — and create new dependencies

The relationship is not one-directional.

AI can help improve load forecasting, equipment monitoring, predictive maintenance, anomaly detection, grid planning, customer service, market analysis, weather interpretation, asset inspection, and operational optimization.

That creates real opportunity.

It also introduces the same governance questions into critical infrastructure.

What data trains the system?

How does it perform during rare events?

Can operators understand and challenge the output?

Could many participants adopt similar models and react in the same way?

What happens if the model, network, or cloud service becomes unavailable?

Can a cyberattack manipulate data or decisions?

Is the model’s recommendation robust outside historical conditions?

Does automation reduce the operator’s ability to recognize an abnormal state?

The correct response is not to reject AI.

It is to use it with an operating model that recognizes the consequences of dependency.

Energy markets reveal the central truth of connected risk: strategy operates inside physical limits. Those limits can be expanded, shifted, contracted around, diversified, or adapted to — but they cannot be governed away by optimistic language.

5. The governance problem is fragmented ownership

Most organizations do not suffer from a complete absence of ownership.

They suffer from ownership that stops at functional boundaries.

Technology owns AI implementation.

Data teams own pipelines and quality.

Cybersecurity owns security controls.

Legal and compliance own policy interpretation.

Sustainability owns climate analysis and reporting.

Facilities owns buildings and backup systems.

Operations owns continuity.

Procurement owns supplier contracts.

Finance or treasury owns market and liquidity exposure.

Business units own performance.

Risk owns frameworks, challenge, aggregation, and reporting.

Each function sees a legitimate part of the picture.

The interaction remains difficult to own.

Who owns the relationship between AI adoption and energy dependency?

Who owns the relationship between a climate hazard and cloud-service continuity?

Who decides whether model concentration is acceptable when switching costs are high?

Who owns the risk that a manual fallback can process only a fraction of normal volume?

Who can delay an AI rollout because infrastructure readiness is weak?

Who decides which workloads receive priority during capacity constraint?

Who accepts the residual risk when the commercial deadline arrives before the resilience control?

Who tells the board that three supposedly separate risks share one underlying dependency?

These are governance questions before they are technical questions.

Shared ownership can become unowned interaction

“Shared ownership” often sounds mature. In practice, it can mean that every team assumes another team is managing the connection.

Technology may believe power resilience is the vendor’s responsibility.

The vendor may rely on its utility and contractual force-majeure terms.

Procurement may have confirmed the service-level agreement.

Legal may have established remedies.

Operations may have documented a workaround.

The business may assume the AI capability is part of normal capacity.

Risk may see four separate controls and conclude that the exposure is covered.

Then a disruption reveals that:

  • The service-level remedy is financial, not operational.
  • The workaround supports only a small percentage of volume.
  • The alternate vendor requires weeks of configuration.
  • The alternate cloud region shares a common dependency.
  • The business has not prioritized customers or services.
  • No executive has explicit authority to suspend lower-value workloads.
  • Communications cannot explain the failure without creating legal or reputational problems.
  • The board has never approved the actual residual risk.

Every team completed its assigned task.

The organization still failed to govern the system.

The risk function should own the connected view

Risk should not try to operate the grid, configure the model, negotiate every contract, run the data center, or own the sustainability program.

It should ensure that leadership receives a connected view of exposure and that the operating functions have made the required decisions.

That means the risk function should challenge across at least five dimensions.

Completeness: Have all material dependencies and consequences been identified?

Coherence: Do the assumptions made by different functions agree?

Concentration: Do nominally separate controls depend on the same vendor, location, infrastructure, person, or decision?

Execution: Can the stated fallback actually operate at the required scale and speed?

Accountability: Is there a named executive who can accept, reduce, transfer, or stop the risk?

This is not administrative aggregation. It is systems challenge.

A strong risk function can say:

“The AI team has approved the model, but operations has not validated the minimum service if capacity is reduced.”

“The climate assessment identifies regional heat risk, but the infrastructure plan assumes historical peak demand.”

“The procurement team has an alternate vendor, but migration has never been tested.”

“The continuity plan assumes staff availability that conflicts with the severe-weather scenario.”

“The market analysis covers average cost but not peak-volume and collateral exposure.”

“The board has approved the strategy but has not approved the concentration and fallback assumptions required to deliver it.”

That is useful risk leadership.

Decision rights matter more than committee structure

Organizations often respond to emerging risk by creating a committee.

Committees can improve coordination. They can also diffuse responsibility.

The important question is not how many people attend. It is what decisions the body is authorized to make.

A connected-risk governance structure should clarify:

  • Who classifies an AI-enabled service as critical?
  • Who sets its impact tolerance?
  • Who approves the dependency and concentration profile?
  • Who can require additional resilience before launch?
  • Who monitors infrastructure and climate indicators?
  • Who declares degraded operation?
  • Who can shift, throttle, or suspend workloads?
  • Who prioritizes customers and services?
  • Who authorizes emergency expenditure?
  • Who communicates with regulators, customers, employees, and the board?
  • Who accepts residual risk, and for how long?
  • Who verifies that remedial action was completed?

A committee that cannot answer those questions may improve awareness without improving control.

Residual risk acceptance must be specific

Residual risk is often described in broad language: “Management accepts the risk.”

That is not enough.

A credible acceptance should identify:

The specific exposure.

The affected service.

The likely and severe consequences.

The control gap.

The reason the gap remains.

The duration of acceptance.

The indicators that would trigger reconsideration.

The person with authority to accept it.

The funded remediation plan.

The conditions under which the activity would be reduced or stopped.

This matters because connected risks often involve commercial pressure.

The business case is approved.

The launch date is public.

The vendor has been selected.

The budget cycle is closing.

The infrastructure solution will arrive later.

At that moment, governance is tested.

A vague statement allows organizational momentum to become the decision-maker.

A specific acceptance forces leadership to acknowledge the trade-off.

Board oversight must cross committee boundaries

AI may be discussed by a technology committee.

Climate may be discussed by a sustainability or risk committee.

Infrastructure may be discussed through capital expenditure.

Energy exposure may be covered in finance.

Operational resilience may sit with audit or risk.

That structure can work only if someone assembles the whole chain.

The board does not need every technical detail. It does need to know whether the organization’s strategic commitments rely on assumptions that have not been tested together.

The central board question is:

What must remain true — physically, digitally, financially, and organizationally — for this strategy to work?

Management should be able to answer with evidence rather than reassurance.

6. From risk register to operating model

A risk register is useful for inventory, accountability, and reporting.

It is not an operating model.

A register might include entries such as:

AI-service outage.

Extreme-weather disruption.

Energy-price volatility.

Cloud-provider concentration.

Data-center capacity.

Supply-chain interruption.

Those entries can be scored, assigned, mitigated, and reported.

What the register usually does not show is sequence.

Which event affects which dependency?

Which control fails first?

Which service degrades?

How quickly does the consequence become material?

Which indicator provides warning?

Which decision must be made?

Who is waiting for whom?

What happens if two risks occur together?

What action is no longer available after a delay?

Connected risk is fundamentally dynamic. It requires a view of how the organization behaves under stress.

Start with the service, not the risk category

Operational resilience frameworks increasingly emphasize critical operations, interdependencies, continuity, and testing. The Basel Committee’s principles for operational resilience, developed for banks but instructive more broadly, focus on the ability to withstand severe disruptions rather than trying to prevent every event.

That is the right starting point.

Choose a critical service.

Define the outcome it must deliver.

Identify the maximum tolerable disruption or degradation.

Map everything required to deliver it.

Then test the system.

For example, consider an AI-enabled customer-support service.

It may depend on:

  • Customer identity systems
  • Telecommunications
  • Cloud infrastructure
  • One or more foundation models
  • Retrieval databases
  • Customer records
  • Content controls
  • Human supervisors
  • Escalation teams
  • Quality monitoring
  • Network access
  • Regional data centers
  • Electricity and cooling
  • Vendor support
  • Payment systems
  • Regulatory notices
  • Employee availability
  • Crisis communications

An “AI outage” is only one failure mode.

The service could also fail because identity is unavailable, customer data is stale, the model changes behavior, response latency increases, demand exceeds contracted capacity, employees cannot reach the fallback site, telecommunications fail, or an extreme event affects several vendors simultaneously.

The operating model must manage the service across those conditions.

Define impact tolerance, not just recovery aspiration

Organizations commonly use recovery-time and recovery-point objectives.

Those are valuable, but connected risk requires a broader impact tolerance.

How long can the service be unavailable?

How long can it operate in degraded mode?

What volume can be deferred?

What error rate is acceptable?

Which customers or transactions must be protected first?

What safety, legal, or conduct boundary cannot be crossed?

At what point does delay become irreversible harm?

What is the maximum financial exposure?

What public commitment has been made?

What dependency has the shortest tolerance?

An organization may have a four-hour technology recovery target while customer harm becomes material after thirty minutes.

It may restore the application in two hours while the data required by the application remains unavailable for a day.

It may recover the AI service but have insufficient staff to process the backlog.

It may meet a technical target while breaching the actual service tolerance.

Impact tolerance should therefore be defined at service level, supported by technology objectives rather than replaced by them.

Risk appetite must be executable

Risk-appetite statements often use language such as “low appetite for disruption to critical services.”

That expresses intent. It does not guide a decision.

An executable appetite includes thresholds.

For example:

No critical customer decision will rely solely on an unvalidated model output.

No critical AI service will enter production without a tested degraded mode.

No single provider will support more than a defined proportion of services above a certain criticality unless an executive accepts the concentration.

No facility or vendor supporting a critical process will be treated as resilient based solely on contractual assurance.

No manual fallback will be credited above the capacity demonstrated in a test.

No infrastructure dependency with a lead time beyond the strategic launch date will remain outside executive reporting.

No residual control gap will be accepted indefinitely.

The exact thresholds vary by organization. The principle does not.

Risk appetite becomes real when it changes behavior before an incident.

Distinguish friction from fragility

Every operating system contains friction.

A process slows.

A vendor has an incident.

A forecast is wrong.

A project is delayed.

A cost rises.

Friction can be inconvenient without threatening the system.

Fragility is different.

Fragility exists when a relatively modest disturbance creates disproportionate consequences because dependencies are concentrated, buffers are thin, fallback is untested, or decisions are delayed.

A resilient organization absorbs friction.

A fragile organization appears efficient until conditions move outside a narrow range.

Risk leaders must help executives tell the difference.

Some useful indicators of structural fragility include:

Rapid growth in critical workloads without corresponding capacity.

Increasing dependence on a small number of providers.

Manual processes that exist only on paper.

Recovery targets supported by assumptions rather than tests.

High utilization with little operational buffer.

Several controls relying on the same people or infrastructure.

Temporary exceptions that become permanent.

Climate adaptation deferred across successive capital cycles.

Market exposure measured only under average conditions.

A strategy whose economics fail under a small change in energy or vendor cost.

An incident process that requires consensus before anyone can act.

These are not isolated control weaknesses. They are signs that the operating model may fail nonlinearly.

Ask what the organization will stop doing

Most continuity plans focus on restoring activity.

Severe constraints also require stopping activity.

During a capacity shortage, not every workload can remain a priority.

During a prolonged outage, not every customer can receive normal service.

During a liquidity event, not every expenditure can proceed.

During a regional disruption, not every facility can be restored at once.

During degraded AI operation, not every use case should remain active.

The organization should decide in advance:

Which services are essential?

Which can be reduced?

Which can be delayed?

Which create more risk than value in degraded conditions?

Which customers or communities require protection?

Which automated decisions should revert to manual review?

Which data processing can be suspended?

Which AI workloads can be shifted to off-peak periods?

Which projects should yield capacity to current operations?

Who authorizes those choices?

This is uncomfortable because it exposes strategic priorities.

That is why it belongs in governance before a crisis.

A connected stress scenario

Consider a plausible scenario.

A prolonged regional heatwave raises electricity demand. Wildfire conditions affect a major transmission route and create air-quality restrictions. The grid remains operating, but reserve margins tighten and market prices rise.

A cloud provider experiences capacity pressure in the affected region. It does not fail completely. Instead, response times increase and some non-guaranteed AI capacity is limited.

At the same time, customer demand rises because the weather has disrupted services across the region. The organization’s AI-enabled support channel begins to queue requests. Automated summaries are delayed. Human agents receive incomplete context. Error rates increase.

The formal alternate cloud region is available, but migration requires a configuration change, additional approval, and data replication that has not been tested at current volume.

The manual procedure can handle 25 percent of normal demand. Many trained employees are dealing with heat, transport disruption, childcare, or air-quality conditions. The backup site is located inside the same regional event.

A key supplier requests accelerated payment because its own energy and logistics costs have risen. The organization’s insurance coverage includes physical damage but provides limited protection for a service interruption originating with a cloud provider.

No single event is catastrophic.

Together, they create a governance crisis.

Does management restrict lower-priority AI workloads?

Does it shift customer traffic?

Who accepts the reduced service quality?

When does the incident become reportable?

Who approves emergency cloud expenditure?

Does the business continue automated decisions under degraded performance?

Which customers receive priority?

What public explanation is accurate?

How long can the organization operate before backlog creates unacceptable harm?

A risk register can contain every component of this scenario and still fail to answer those questions.

An operating model must answer them.

7. A practical framework for connected risk governance

Organizations do not need another abstract taxonomy.

They need a repeatable method for connecting strategy, dependencies, thresholds, controls, and decisions.

The following eight-part framework can be applied to a business service, major AI program, infrastructure investment, critical supplier, facility portfolio, or strategic transformation.

The objective is not to predict every event.

It is to ensure that the organization understands what it depends on, how those dependencies can fail, and what leadership will do when conditions move outside plan.

1. Build the dependency map

Begin with a critical service or strategic outcome.

Do not begin with the organization chart.

Ask what must work for the service to be delivered from end to end.

The map should include:

Business processes.

AI models and applications.

Data sources and pipelines.

Cloud regions and availability zones.

Networks and identity systems.

Physical facilities.

Electricity, cooling, water, and telecommunications.

Hardware and specialist equipment.

Employees and contractors.

Critical suppliers and subcontractors.

Legal permissions and regulatory conditions.

Financial capacity and payment arrangements.

Customer and community dependencies.

Emergency communications.

The map should show both upstream and downstream consequences.

Upstream mapping identifies what the organization requires.

Downstream mapping identifies who is affected if the service fails.

The most valuable part of the exercise is often the identification of common dependencies.

Several applications may rely on one identity platform.

Several vendors may use the same cloud region.

Two backup facilities may depend on one transmission corridor.

Multiple manual procedures may require the same small group of employees.

An alternate supplier may depend on the same port, substation, telecommunications route, or specialist component.

A dependency map should therefore answer four questions:

What is required?

Where is it located?

Who controls it?

What else fails with it?

Practical questions for risk leaders include:

  • Which dependencies are not visible in current architecture or supplier records?
  • Which critical components are fourth-party or fifth-party services?
  • Where does apparent diversity conceal common infrastructure?
  • Which dependency has the longest replacement lead time?
  • Which component has the shortest failure tolerance?
  • Which resource operates closest to capacity?
  • What information would be needed to switch providers?
  • Which dependencies are outside contractual audit rights?
  • Which people possess knowledge that is not adequately documented?
  • Which manual process depends on technology that may fail in the same incident?

The output should not be a diagram so complex that only its creator can read it.

A useful dependency map is layered.

The board may need a one-page view showing the critical service and major concentrations.

Executives need thresholds, accountable owners, and key alternatives.

Operating teams need the detailed architecture, procedures, contacts, and technical dependencies.

The map should be updated when the service changes. AI programs evolve quickly; a dependency map approved at launch can become obsolete within months.

2. Create a climate and infrastructure exposure map

The second step is to overlay physical exposure.

This should be more specific than a national or regional climate score.

Critical dependencies must be connected to actual locations and operating conditions.

For each material site, supplier, cloud region, network route, facility, and infrastructure node, assess:

Relevant acute hazards.

Relevant chronic changes.

Current protective controls.

Design assumptions.

Maintenance condition.

Historical incidents.

Future operating life.

Recovery capability.

Alternative capacity.

Insurance and financial protection.

Dependency on local public infrastructure.

The assessment should examine multiple hazards, not one at a time.

A facility protected against flood may remain vulnerable to grid outage.

A data center with backup power may depend on water, fuel delivery, or telecommunications.

A supplier outside the flood zone may rely on a transport route inside it.

A cloud region may be geographically separate but exposed to the same heat or drought pattern.

The risk function should challenge false precision.

Climate models and hazard data have uncertainty. So do demand forecasts, asset-condition data, and supplier disclosures.

Uncertainty is not a reason to ignore exposure. It is a reason to use ranges, scenarios, and decision triggers.

Ask:

  • What conditions were used when the asset or control was designed?
  • Have those conditions changed?
  • What is the difference between the expected case and severe-but-plausible case?
  • Which exposure is already close to an operating threshold?
  • Which adaptation action requires the longest lead time?
  • What happens if two hazards occur together?
  • Does the alternate site face the same regional event?
  • Which local public services are assumed to remain available?
  • Are community and workforce impacts included?
  • Which suppliers lack the financial capacity to recover?
  • Where is insurance being treated as a substitute for operational resilience?
  • Which assets may remain operable but become uneconomic?

The output should identify priority adaptation actions and their decision dates.

A control that must be completed in four years should not be governed as though management can decide in year four. Engineering, permitting, procurement, financing, and construction may require action much earlier.

3. Conduct an AI operating-dependency review

Traditional AI review asks whether a model is appropriate, accurate, secure, fair, and compliant.

The operating-dependency review asks what happens to the business when the AI capability changes or disappears.

Classify AI use cases by criticality.

A low-criticality use case might improve drafting or research.

A moderate use case may influence internal analysis but remain subject to meaningful review.

A high-criticality use case may directly affect customers, safety, financial positions, legal obligations, critical infrastructure, or service continuity.

Criticality should determine resilience requirements.

For each material AI-enabled service, review:

The approved model and provider.

Hosting environment.

Data and retrieval dependencies.

Integration architecture.

Identity and access.

Monitoring.

Capacity commitments.

Change rights.

Model-update processes.

Fallback model.

Manual alternative.

Human review capability.

Incident classification.

Exit and portability.

The review should test more than complete outage.

Scenarios should include:

Latency deterioration.

Reduced capacity.

Model-version change.

Loss of retrieval data.

Incorrect or stale source information.

Security restriction.

Vendor-policy change.

Regional cloud failure.

Loss of a specialist integration.

Rapid cost increase.

Regulatory restriction.

Data-sovereignty conflict.

Unexpected demand surge.

Quality degradation that remains above simple availability thresholds.

Ask:

  • What business outcome does the AI system influence?
  • Can the organization detect degraded performance quickly?
  • What output differences occur during failover?
  • Can a human understand why the system made a recommendation?
  • How much volume can humans review?
  • What happens when the exception queue exceeds that capacity?
  • Does the contract guarantee capacity or only access?
  • Can the provider change the underlying model?
  • What notice is required?
  • Has migration been tested?
  • How much historical data, configuration, and evaluation evidence can be exported?
  • What AI capability would the organization suspend first?
  • Is the organization preserving the underlying human skill?
  • What leading indicators show dependency is becoming critical?

The output should be a clear operating profile for each critical service: normal mode, degraded mode, fallback mode, suspension threshold, and decision owner.

4. Review market and liquidity sensitivity

Connected risk can turn a physical or technological problem into a financial problem quickly.

The organization should assess how changes in energy, infrastructure, vendor, insurance, and service conditions affect cash flow, liquidity, collateral, margins, and investment capacity.

This is not limited to companies that trade commodities.

A manufacturer may face higher power and input costs.

A digital company may face cloud-price changes and capacity premiums.

A utility may face load uncertainty, fuel costs, capital requirements, and customer affordability concerns.

A financial institution may face collateral movements and customer stress.

A logistics business may face fuel, transport, and facility disruption.

A supplier may request different payment terms after a regional event.

The review should consider:

Energy volume and price.

Regional basis or congestion.

Peak and shape exposure.

Contract pass-through.

Capacity charges.

Vendor price changes.

Emergency procurement.

Insurance deductibles and exclusions.

Uninsured interruption.

Collateral and margin requirements.

Customer credit deterioration.

Supplier liquidity.

Working-capital needs.

Capital-project delay.

Repair-cost inflation.

Revenue loss during degraded service.

Ask:

  • Which assumptions make the business case work?
  • How sensitive is the plan to energy or infrastructure cost?
  • Which contracts transfer cost but retain operational risk?
  • Which contracts transfer operational responsibility but provide only financial remedies?
  • What cash requirement could arise during a prolonged event?
  • Which counterparties become weaker under the same scenario?
  • Could the organization face higher costs while revenue is also falling?
  • What liquidity is available without impairing recovery investment?
  • Are insurance limits aligned with realistic interruption duration?
  • What happens at renewal?
  • Which resilience investments reduce several risk categories at once?
  • Where do different functions use conflicting price or demand assumptions?

The output should not be one forecast.

It should be a range of sensitivities linked to management actions.

At what point is the program resized?

At what point is a contract renegotiated?

At what point is additional liquidity reserved?

At what point is a supplier replaced?

At what point does infrastructure cost make a use case uneconomic?

These are governance thresholds, not merely financial-model outputs.

5. Review controls and fallback capability

Controls should be examined across the full disruption lifecycle:

Prevent: Reduce the probability of failure.

Detect: Identify deterioration before consequence expands.

Respond: Contain the incident and protect critical service.

Recover: Restore safe and stable operation.

Adapt: Change the system so the same weakness does not persist.

A connected control review should test whether controls are truly independent.

Two cloud regions are not necessarily independent.

Two generators sharing one fuel source are not necessarily independent.

Two suppliers using one subcomponent are not necessarily independent.

Two manual workarounds requiring one person are not independent.

Control design should therefore consider diversity, not only duplication.

For every critical service, define the minimum viable operation.

This is the smallest service the organization must preserve to remain safe, lawful, and operationally coherent.

Then test the fallback against realistic conditions.

How many transactions can it process?

How many employees are required?

What information must be available?

How long can it operate?

What backlog accumulates?

What error rate occurs?

Which customers are prioritized?

What other activities must stop?

How is the transition controlled?

How is normal service restored without data loss or duplicate action?

A fallback that has never been tested at realistic volume should not receive full control credit.

A backup generator that has not been tested under load should not receive full credit.

An alternate model that has not been evaluated on critical use cases should not receive full credit.

A supplier letter stating that continuity arrangements exist should not receive the same credit as evidence from a test.

Risk reporting should distinguish:

Designed control.

Implemented control.

Tested control.

Demonstrated effective control.

Those are not the same state.

6. Establish executive decision rights

Connected risks become dangerous when everyone waits for more information.

Decision rights should be agreed while conditions are normal.

A decision matrix should identify who can:

Declare a connected-risk incident.

Move a service into degraded mode.

Suspend an AI use case.

Prioritize workloads.

Shift customer traffic.

Authorize emergency spending.

Change providers.

Invoke manual processing.

Accept reduced service quality.

Notify the board.

Communicate externally.

Request regulatory relief where available.

Accept temporary residual risk.

End the incident.

Decision authority should be linked to thresholds.

For example:

When AI response latency exceeds a defined level for a defined period, the service owner moves to degraded mode.

When quality monitoring falls outside tolerance, automated decisions stop and manual review begins.

When regional infrastructure indicators reach a trigger, nonessential workloads are reduced.

When manual backlog reaches a threshold, customer prioritization rules activate.

When financial exposure exceeds a limit, treasury joins incident command.

When the control gap remains beyond a defined duration, the accountable executive must renew or terminate the risk acceptance.

The organization should also define who decides when data is incomplete.

A crisis rarely presents perfect information.

The decision process must establish the minimum evidence required, the person authorized to act, and the conditions for reversal.

This prevents consensus from becoming a hidden control weakness.

7. Rehearse compound stress

A tabletop exercise should not be a scripted conversation that ends with reassurance.

It should force decisions.

Use scenarios that combine domains:

Extreme weather plus cloud degradation.

Grid constraint plus customer-demand surge.

AI-model change plus regulatory inquiry.

Vendor outage plus employee unavailability.

Energy-price spike plus supplier liquidity stress.

Flood plus telecommunications failure.

Cyber incident during physical disruption.

Data-quality failure during market volatility.

Introduce events in sequence.

Do not reveal the full scenario in advance.

Require participants to use actual contact lists, procedures, authority levels, dashboards, and fallback tools.

Measure:

Time to detect.

Time to classify.

Time to decide.

Time to communicate.

Time to activate fallback.

Fallback capacity.

Data availability.

Quality under degraded conditions.

Backlog growth.

Financial exposure.

Conflicts in authority.

Assumptions that proved false.

An effective exercise creates controlled discomfort.

It should expose where leaders hesitate, where teams use different data, where contracts fail to answer operational questions, and where the organization cannot determine who is accountable.

Technical tests should complement executive exercises.

Fail over the service.

Run the alternate model.

Operate manually.

Test backup power.

Verify fuel.

Simulate loss of a data source.

Contact the vendor outside business hours.

Restore from backup.

Move a workload.

Reconcile data after recovery.

A rehearsal is valuable only when it changes the operating system.

8. Build a post-stress learning loop

Organizations often conduct an exercise, produce a report, and return to business.

That creates activity without adaptation.

Every incident, near miss, test, vendor event, weather disruption, capacity warning, or market shock should feed a structured learning process.

The review should ask:

What happened?

What did we expect?

Which assumption was wrong?

Which indicator arrived too late?

Which control worked?

Which control existed only on paper?

Where did authority become unclear?

What consequence was avoided by luck?

Which dependency changed without governance noticing?

What should be stopped, redesigned, funded, or accelerated?

Each action needs an owner, deadline, funding source, and verification method.

Lessons should also update strategy.

If AI dependency is growing faster than fallback capacity, the operating model must change.

If climate exposure invalidates a site assumption, the capital plan must change.

If energy constraints alter the economics of a program, the roadmap must change.

If vendor concentration cannot be reduced, risk acceptance and service design must change.

If a control repeatedly fails testing, management should stop describing it as effective.

The board should see persistent themes, not every minor action.

Which dependencies are becoming more concentrated?

Which resilience investments remain delayed?

Which risk acceptances have been renewed?

Which scenarios produce the largest decision gaps?

Where is the organization relying on optimism?

The learning loop is what converts resilience from a project into a capability.

8. Leadership under converging risk

Connected risk is not solved by framework alone.

It is a leadership test.

The technical work can identify dependencies, scenarios, thresholds, and controls. Senior leaders still have to confront uncomfortable trade-offs.

Growth versus resilience.

Speed versus readiness.

Efficiency versus buffer.

Concentration versus cost.

Automation versus human capacity.

Standardization versus diversity.

Short-term performance versus long-term adaptability.

These choices rarely have a perfect answer.

They do require honest assessment.

Discipline is more useful than confidence

Outside the office, I am drawn to disciplines where preparation is exposed quickly: endurance training, motorsport, and tactical preparation.

Each teaches a version of the same lesson.

Ambition does not override operating conditions.

In endurance training, a strong objective does not eliminate heat, hydration, pacing, fueling, fatigue, or injury. A plan that assumes perfect conditions is not aggressive; it is weak.

In motorsport, speed is constrained by grip, tires, fuel, temperature, mechanical condition, track position, and the quality of information reaching the driver and team. Pushing harder can improve performance until it crosses the system’s limits. Beyond that point, more aggression increases the probability of failure.

In tactical preparation, the value of a plan is not how impressive it sounds in a briefing. It is whether people can execute it under uncertainty, degraded communication, incomplete information, and time pressure.

The analogy matters because modern risk leadership faces the same temptation: to treat determination as a substitute for preparation.

It is not.

You do not rise to ambition under pressure. You fall to the operating system you built.

That operating system includes people, habits, data, authority, controls, buffers, and rehearsal.

Honest assessment is a leadership control

Many significant risks remain unresolved because people are rewarded for confidence.

Project sponsors need approval.

Vendors want the contract.

Executives want growth.

Teams want to demonstrate progress.

Risk leaders want to be seen as commercial rather than obstructive.

Those incentives can produce language that is technically defensible but operationally misleading.

“The vendor has high availability.”

“We have an alternate.”

“The process can be completed manually.”

“The climate risk is long term.”

“The cost can be passed through.”

“The model is advisory.”

“The infrastructure is the provider’s responsibility.”

Each statement may contain truth.

The risk leader’s job is to ask what is missing.

High availability under what conditions?

An alternate with what capacity and migration time?

Manual at what volume?

Long term relative to which decision?

Passed through to whom, and with what demand effect?

Advisory in policy or in actual behavior?

Provider responsibility with what operational remedy?

Good challenge is not cynicism. It is precision.

Controlled exposure enables ambition

The purpose of risk leadership is not to remove uncertainty or say no to technology.

The purpose is to prevent the organization from taking exposures it does not understand, cannot monitor, or is unable to manage.

That often produces a conditional yes.

Yes, launch the AI service after the degraded mode is tested.

Yes, use the concentrated provider while the exit capability is built and the board accepts the interim risk.

Yes, expand the facility after power and water assumptions are independently validated.

Yes, proceed with the supplier while alternate inventory and financial triggers are established.

Yes, automate the process while preserving human competence and setting a suspension threshold.

Yes, pursue the strategy while sequencing demand to match infrastructure readiness.

A conditional yes is not weaker than an unconditional yes.

It is more executable.

Calm leadership depends on precommitment

Pressure changes behavior.

People narrow their attention.

Teams protect their own area.

Leaders delay bad news.

Organizations continue familiar activity even when conditions have changed.

Precommitted thresholds reduce that risk.

A marathon plan decides pacing before fatigue distorts judgment.

A race team agrees pit criteria before emotion takes over.

An incident team defines escalation and stop rules before commercial pressure becomes acute.

Connected-risk governance should do the same.

Decide in advance what conditions trigger degraded operation.

Decide which services receive priority.

Decide when automation stops.

Decide who can act without full consensus.

Decide what information reaches the board.

Decide which residual risks expire automatically.

Precommitment does not remove judgment. It protects judgment from predictable pressure.

Risk credibility is earned through operational usefulness

A risk function will not earn influence by producing the most categories.

It earns influence when it helps the organization make difficult decisions earlier and better.

That requires risk leaders to understand the business model, operating architecture, markets, infrastructure, and human realities well enough to challenge with specificity.

The risk leader should be able to move between the boardroom and the operating detail.

At board level:

What strategic assumption is most vulnerable?

At executive level:

Which decision right is missing?

At operating level:

Which dependency, threshold, or control is untested?

The future chief risk officer is not simply the owner of a framework.

The role is becoming that of a systems integrator: someone who can connect technology, climate, infrastructure, markets, operations, capital, and accountability without pretending to replace the experts in any one of them.

That is a demanding standard.

It is also the standard the operating environment now requires.

9. What boards and executives should ask

Boards do not need to become AI engineers, climate scientists, or power-market operators.

They do need to ask questions that force management to connect strategy with evidence.

The most useful questions are those that reveal assumptions, concentrations, thresholds, and authority.

Questions about strategy and dependency

Where are we treating infrastructure as an assumption rather than a governed dependency?

This includes power, cloud capacity, telecommunications, water, transport, facilities, specialist hardware, and people.

Which parts of our AI strategy depend on physical capacity that is not yet secured?

Management should distinguish contracted access, expected access, proposed infrastructure, and demonstrated capacity.

Which AI capabilities have moved from optional tools to critical operating dependencies?

The answer should be based on actual workflow and service commitments, not the original project classification.

What must remain true for the business case to work?

This should include cost, capacity, vendor performance, energy availability, workforce productivity, regulatory permission, and customer adoption.

Questions about climate and infrastructure

Which climate scenarios would change our operating model rather than only our disclosures?

Management should identify affected services, thresholds, controls, investments, and decisions.

Where do several critical dependencies share one geographic exposure?

Cloud regions, suppliers, offices, logistics routes, utilities, telecommunications, and employee populations should be considered together.

Which assets or facilities are operating against historical assumptions that may no longer be adequate?

The answer should address design standards, maintenance, insurance, operating life, and adaptation plans.

Which resilience investments are being deferred because the benefit falls across several business units?

Cross-functional benefits often create budget gaps.

Questions about concentration and fallback

Which risks are owned across several teams but mastered by none?

Management should be able to name the executive responsible for the interaction.

Where does apparent vendor diversification conceal a common underlying provider or infrastructure dependency?

The board should expect visibility beyond first-tier contracts.

What proportion of normal activity can our fallback actually support?

The answer should come from a test.

When did we last operate a critical service without its primary AI, cloud, power, data, or telecommunications dependency?

A plan that has not been exercised is an assumption.

What capability have we allowed to atrophy because automation is working?

This may include human expertise, manual processing, local technical skill, supplier diversity, or spare capacity.

Questions about decision rights

Who can declare degraded operation?

Who can suspend a critical AI use case?

Who decides which workloads or customers receive priority under constraint?

Where do we require consensus when one accountable executive should act?

Which residual risks have been accepted, by whom, and until when?

The board should be cautious when decision authority is described only through committee membership.

Questions about economics and market sensitivity

How does the strategy perform under higher energy, infrastructure, cloud, insurance, or supplier costs?

Which contracts protect price but not availability?

Which contracts provide compensation after failure but no viable operational replacement?

Could a single event increase cost, reduce revenue, and create liquidity demand simultaneously?

What financial capacity has been reserved for recovery and adaptation?

Resilience is weakened when emergency response competes with ordinary liquidity at the worst moment.

Questions about learning

What did our last stress exercise cause us to change?

If the answer is only documentation, the exercise may not have been demanding enough.

Which control repeatedly appears in reports but remains untested?

What near miss are we treating as good luck rather than evidence of weakness?

Which dependency is growing faster than our understanding of it?

What would we stop doing first if constraints tightened tomorrow?

That final question is particularly important.

It reveals whether management has genuinely prioritized the operating model or simply labeled everything critical.

Boards should not seek certainty from these questions.

They should seek disciplined evidence that management understands the conditions under which strategy succeeds, the conditions under which it degrades, and the decisions required when those conditions change.

10. The seat risk leadership must earn

AI risk will not remain inside the technology function.

Climate risk will not remain inside sustainability reporting.

Infrastructure risk will not remain buried in a vendor contract.

Energy-market risk will not remain a finance or trading issue.

Operational resilience will not be protected by a continuity document alone.

These risks are becoming connected because the operating model is becoming connected.

AI expands digital dependency.

Climate conditions stress physical systems.

Energy sits between demand and capacity.

Markets transmit scarcity.

Third parties concentrate infrastructure.

Automation changes the role and availability of people.

Executive decisions determine whether pressure becomes controlled degradation or uncontrolled failure.

The organization can connect these realities in governance, or reality will connect them during a disruption.

That is the choice.

A mature organization does not need perfect forecasts. It needs clear dependencies, credible scenarios, tested controls, executable thresholds, and leaders who know what authority they hold.

It needs to understand the difference between a contractual remedy and an operational alternative.

Between a backup and a demonstrated fallback.

Between a risk score and an impact tolerance.

Between nominal diversification and physical diversity.

Between sustainability reporting and adaptation.

Between AI access and AI fitness for purpose.

Between average energy cost and capacity during stress.

Between confidence and preparation.

The role of risk leadership is not to oppose AI, climate strategy, infrastructure investment, or growth.

It is to make ambition operable.

That means showing where the strategy relies on physical capacity.

It means challenging assumptions that cross functional boundaries.

It means insisting that residual risk has a name, an owner, a duration, and a stopping rule.

It means helping management preserve options before those options become expensive or unavailable.

It means rehearsing the decision, not just documenting the procedure.

It means maintaining enough humility to recognize that efficient systems can become fragile when buffers disappear and dependencies concentrate.

The risk function earns its seat when it can help the organization move forward without confusing momentum with resilience.

The future risk leader will not be defined by the ability to say no to technology, climate commitments, infrastructure investment, or growth.

The future risk leader will be defined by the ability to connect ambition to operating reality — before reality does it by force.


메타데이터
post_id
2d1a18f005fa
slug
ai-climate-and-infrastructure-risk-the-new-governance-test-2d1a18f005fa
url
https://medium.com/@kevin.pausicles/ai-climate-and-infrastructure-risk-the-new-governance-test-2d1a18f005fa
canonical_url
https://medium.com/@kevin.pausicles/ai-climate-and-infrastructure-risk-the-new-governance-test-2d1a18f005fa
author_url
https://medium.com/@kevin.pausicles
status
ok
fetched_at
2026-08-23 12:25:54