← Back to list

When the AI Was Right, and the Optimization Still Went Wrong

Why autonomous energy system failures rarely come from the algorithm, and what designers must do differently

Hamed Sattarian · 2026-02-27 20:22 · 0 claps · 17.1 min read
#energy-consumption #ai-ux-design #ai-product-design #ai-safety #hax
Open on Medium ↗
Wiki topics: AGT · AI Agents SAF · Safety & Alignment UX · UI/UX Design PRD · Product Design DSN · Design · General 💻 · Programming

When the AI Was Right, and the Optimization Still Went Wrong

Why autonomous energy system failures rarely come from the algorithm, and what designers must do differently

A note: The McDonald’s scenario in this article is entirely fictional and not based on any real incident. It’s a composite I built from failure patterns encountered while working on a different energy optimization product, reframed here as a hypothetical for illustrative purposes.

Imagine a McDonald’s franchise in suburban Chicago, early February morning. This location is one of 1,400 that signed on to an AI-powered energy optimization platform about a year and a half ago. And based on the dashboard metrics, it’s been a big win: HVAC energy costs are down 22%, cooling expenses have dropped 19%, and they saw a return on investment by month eleven. Those are the kind of figures that make it into the quarterly business review slides.

At 6:47 AM, the outside temperature begins to fall faster than expected. The AI kicks in, just like it has every morning, pre-cooling the dining area in preparation for the breakfast crowd, trying to strike a balance between comfort and peak demand costs. What the AI doesn’t realize — and what the dashboard isn’t showing — is that the rooftop temperature sensor has been off for three weeks, reading consistently 4°F higher than it should. Because of this drift, the system has been working harder than necessary, running the compressors more often, and all the overall metrics remain in the green, since green is the only color the dashboard displays.

By 8:15 AM, the dining area is a chilly 61°F. The breakfast manager presses the “manual override” button, believing she’s taking charge of the situation. But she’s not fully in control. The override stops setpoint recommendations, but compressor commands continue as they are. By 9:30, a technician is on the scene. And by 11:00, two compressors have failed due to excessive short-cycling.

The algorithm didn’t mess up; the sensor drift is just a common, well-known issue in commercial HVAC systems. The real failure happened in the interface — three specific areas where things went wrong — before anyone in the building even realized there was a problem.

This is the kind of issue that really needs more straightforward focus in building automation design. It’s not just about whether the AI works well under perfect conditions; it’s about whether the entire system — including everything the operator sees — is ready for when something goes off track.

The dashboard said everything was fine. The building was saying something else entirely.

The dashboard said everything was fine. The building was saying something else entirely.

The trust problem begins in the design room, not the control room

When AI products fail in the real world, it’s all too common for post-failure discussions to zero in on the operator. Questions like, “Why didn’t she check the sensor logs?” or “Why did she just assume the override was functional?” tend to pop up. It’s also common to hear, “Why wasn’t she better acquainted with what ’manual mode’ actually entails in this system?”

But these aren’t the right questions. In fact, they’re questions that come up too late. The groundwork for the Chicago failure was laid long before that fateful Tuesday morning, rooted in a series of product decisions that, while reasonable on their own, led to an interface that left operators out of sync with reality.

In 2004, researchers John Lee and Katrina See explored how trust develops in automation and identified three main channels: what the system does (Is it reliable?), how it works (Can an operator follow its thought process?), and why it exists (Are the system’s objectives in line with the operator’s goals?). Most AI product designers focus heavily on the first channel. They present energy dashboards, accuracy metrics, and uptime stats — all proof that the system has been correct. Yet this gives no insight into when it might fail or how an operator can identify that.

In Chicago, the franchise manager had a year and a half of trust built on consistent performance. That’s quite a bit of reassurance. But it fostered trust in average outcomes, whereas what was truly needed was a more nuanced trust calibration — a real-time understanding of whether the current inputs were trustworthy, if the system’s confidence was justified, and whether it was safe for the system to operate autonomously or if a human needed to step in and pay closer attention.

Raja Parasuraman, a veteran in studying human interactions with automated systems, laid out three types of miscalibrated trust. First, there’s misuse: operators rely on automation even when it’s not the right choice, because it’s been accurate so often that questioning it feels pointless. Then there’s disuse: operators dismiss automation that is, in fact, correct, usually because it has made mistakes previously and lost its credibility. Lastly, there’s abuse, which typically happens during product development rather than in operational settings: designers and engineers automate features without fully considering the implications for the users down the line.

The Chicago incident falls into the abuse category. It wasn’t out of malice, there was no intention to keep sensor health data hidden from operators. But at some point during product development, a decision was made to treat sensor status as a maintenance issue rather than something operators should monitor. That choice, along with others like it, led to three interface failures on that Tuesday morning. And each of those failures stemmed from design decisions rather than operator mistakes. Three failures in sequence: a close reading

Breaking down what actually happened between 6:47 and 11:00 AM in that Chicago location reveals how interface failures compound.

The first failure: the sensor never showed up. Sensor drift isn’t some rare occurrence in commercial HVAC systems; it’s pretty standard and well-known. Sensors gradually drift out of calibration, and the AI tries to compensate by making other equipment work harder. This keeps the performance metrics looking good, even as the system itself is falling apart. Lisanne Bainbridge pointed this out way back in 1983, calling it “automation camouflage” — the control system conceals issues by managing around them, until it can’t anymore.

In a well-designed user interface, sensor health should be a top priority — it needs to be right there on the main operating screen, next to every metric that relies on that sensor’s data. Not buried in some maintenance submenu that only technicians can reach. The drift from the rooftop unit should have been displayed alongside the temperature reading it was providing, with a simple indicator saying, “sensor variance: elevated — 3 weeks outside normal range.” That’s not rocket science. It’s really just a product decision to ensure this info is visible to operators rather than hidden away in a log.

The second failure: the override did something completely different than the operator expected. When the manager switched to “manual mode,” she had a clear idea in her mind: I’m taking control, the AI is stopping. That’s what “manual” means in everyday terms. But in the way the product was built, “manual mode” only paused the setpoint suggestions, while the core compressor control logic kept working. The disconnect between what the interface suggested would happen and what actually occurred was where things fell apart.

This is what we call mode confusion — it’s a similar issue that played a role in the Air France 447 incident back in 2009, where pilots who turned off the autopilot didn’t realize that other automation layers were still in play and reacted to a stall by fully pulling back on the stick, thinking they had full control, because the interface didn’t provide a clear picture of the aircraft’s actual status. The cockpit voice recorder caught the moment they realized the reality: “I’ve had the stick back the whole time!” Three and a half minutes into a situation that could have been recovered, the crew was essentially flying blind. The Chicago manager’s scenario wasn’t as serious or tragic, but it was fundamentally the same: an interface that allowed for real misunderstandings about what was under human control.

A safety hierarchy to stop these kinds of mistakes is clear, consistently labeled, and always visible. The system should be in one of four states: full autonomous operation, supervised autonomy within pre-set boundaries, schedule mode that follows defined rules, or emergency fallback that holds the last stable setpoints. The current state should be front and center on every screen. If “manual mode” isn’t fully manual, it’s not a real mode — it’s just asking for trouble.

The third failure: no one could figure out what the AI had done or why. When the technician got there at 11:00 AM, his first question was: what did the system do between 6:47 and 9:30? In most cases, the honest answer would be: there’s a log somewhere, in a format that engineers can read, showing the commands with timestamps. But no interface component lets the franchise manager see, in straightforward terms: “At 7:12 AM the AI bumped the compressor output to 84% because the rooftop sensor reported 41°F and the forecast predicted a warming trend — but the actual outdoor temperature at that time was only 37°F.”

Without that reconstruction, the manager can’t understand what happened. The technician can’t diagnose efficiently. The product team learns nothing about how their AI behaves under degraded input conditions. And the next time a sensor starts drifting, no one has the mental model to recognize the early pattern.

Sensor drift. Ambiguous override. Empty post-incident screen. Three decisions made during product development.

Sensor drift. Ambiguous override. Empty post-incident screen. Three decisions made during product development.

What explainability actually requires when you’re managing fifteen locations

There is a recurring gap in building AI product development between what the engineering team means by “explainability” and what an operator actually needs to make a good decision.

The explainability tools commonly used in AI, like SHAP and LIME, focus on pinpointing which input factors influenced a model’s output. However, their visualisations tend to be technically correct but practically unhelpful in many real-world situations. For instance, if a franchise manager learns that “outdoor temperature was the most important factor in this decision,” it’s not really actionable for her. She can’t validate whether the reasoning is accurate, can’t challenge it, and can’t even tell if an input affecting that “most important factor” is flawed.

IBM researcher Q. Vera Liao spent years figuring out the actual questions non-expert users ask when seeking to understand AI decisions. Interestingly, the most frequent type of question isn’t “why did this happen?” but rather contrastive ones like “why this and not something else I expected?” The second most common type is counterfactual, asking questions like “what would need to change for the outcome to be different?” These questions are critical for someone who needs to decide whether to trust or disregard a recommendation.

Take the Chicago scenario as an example: at 7:00 AM, the franchise manager didn’t need a feature importance chart. What she really needed was a clear, straightforward line: “System is working harder than usual for this temperature. Cross-check indicates the rooftop sensor might be reading higher than actual conditions. Best to verify the sensor before the rush.” This line addresses a contrastive question by explaining why the system is operating more than expected, highlighting an issue with the input, and gives her a specific action to take.

Creating interfaces that tackle these questions necessitates a design framework that focuses on what different users need at various moments, rather than what the engineering team finds intriguing to show off.

A progressive disclosure model that fits multi-site franchise operations can be structured in three layers. The main view — the first thing the manager sees when she opens the app — provides a clear statement of the current AI operating mode, a brief summary of what it’s doing and why, as well as any anomalies that need attention before the next shift. Diving a level deeper, it shows the top three factors influencing current behaviour in simple terms (for example, “outdoor temperature forecast: dropping, which is causing increased pre-cooling”). And at the deepest level, there’s a counterfactual view that states: “if the rooftop sensor were reading normally, the AI would be operating at 60% capacity instead of 84%.”

The manager doesn’t need to grasp reinforcement learning to understand that third layer; she just needs to recognise that something’s not right and know how to address it.

Another useful feature that most platforms miss is per-decision confidence indicators. Telling a franchise manager that “the system is 94% accurate” doesn’t help much on a morning when a specific input sensor has been misbehaving for weeks. However, saying “current recommendation: low confidence — temperature input is outside the normal variance range for this sensor” gives her the vital signal she needs to either resolve an issue at 7:00 AM or end up calling a technician at 11:00.

The first layer is for a normal Tuesday. The third layer is for when something doesn’t feel right.

The first layer is for a normal Tuesday. The third layer is for when something doesn’t feel right.

Designing for the 5% — the cases that actually determine trust

Most interface design for autonomous systems is built around the 95% case: the AI is working correctly, inputs are reliable, the operator’s role is to monitor and occasionally review. This is understandable. The 95% case is what gets demonstrated in sales cycles, piloted in controlled deployments, and shown in product review sessions.

The 5% case — sensor degradation, model drift after occupancy pattern changes, compressor stress from aggressive optimization, edge weather conditions outside the training distribution — is where trust is actually tested and where most products are least prepared.

Bainbridge’s insight from 1983 remains the most uncomfortable truth in this space: the more reliable the automation, the less capable the humans operating it become at taking over when it fails. Eighteen months of successful operation in Chicago meant eighteen months of not needing to understand the HVAC system at a mechanical level. That’s not negligence. That’s how human skill atrophy works. By the time the manager needed to act as a genuine backup operator, the interface and the deployment model had not maintained her in that role.

For a product serving franchise locations, five failure modes deserve explicit design attention — not as edge cases to be handled in documentation, but as first-class scenarios that inform interface architecture.

Compressor short-cycling. AI optimization pursuing tight setpoints can command rapid on/off cycling if minimum runtime constraints aren’t enforced at the control layer. Compressors are rated for a limited number of starts per hour; exceeding that causes winding degradation that doesn’t appear on an energy dashboard. The interface should surface when the AI’s actions are approaching mechanical stress thresholds, not just whether the efficiency target was hit.

Sensor drift under compensation. As described above: the AI compensates for a bad reading by working harder, aggregate metrics stay stable, the problem hides. Every metric the AI is acting on should have a health indicator attached. On the main screen. Not in a maintenance log.

Model drift after pattern changes. The Chicago deployment was trained on data from before a significant shift in dining patterns — delivery orders now represent a higher share of activity, the kitchen load profile has changed, the in-person lunch peak is different. A model trained on the old pattern optimizes for loads that no longer exist at the times they no longer exist. The interface should flag when current operating conditions diverge significantly from the model’s training distribution. Operators can’t act on drift they can’t see.

Override scope ambiguity. “Manual mode” that isn’t fully manual. “Pause” that stops recommendations but not commands. “Override” that affects one subsystem but leaves others running. Every intervention in the interface should carry an explicit scope statement. If that statement would confuse an operator, the underlying feature needs to be redesigned.

Graceful handoff when AI exits. When the system returns control to a human, it should brief the operator rather than simply step aside. Current state of all relevant setpoints. What the AI was attempting to achieve. What condition triggered the exit. Recommended first manual action. An AI that transfers control without context is not providing a safety feature — it’s transferring an unmanaged situation to someone who wasn’t tracking it.

The fallback hierarchy itself — full autonomous, supervised autonomy, schedule mode, emergency fallback — should be visible on every screen, at all times, without requiring navigation. Not as a status badge in a corner. As the primary operational context for everything else being displayed.

How the McDonald’s deployment should have been structured

The failure in Chicago was not the inevitable outcome of deploying building AI in a franchise context. It was the predictable outcome of a deployment model that skipped the stages where trust gets built.

The pattern that produced the most successful large-scale building AI deployments — DeepMind’s handover of Google’s data centers, Phaidra’s staged rollout at Merck’s pharmaceutical manufacturing campus — shares a common architecture: autonomy is earned incrementally, through demonstrated performance, at a pace operators can verify and understand.

In the Google data center work, the transition to full autonomous control happened because the human operators who had been reviewing AI recommendations for months concluded, on their own, that they trusted the system enough to stop reviewing every decision. The engineers didn’t impose Phase 3. The operators asked for it. That distinction — between autonomy granted from above and autonomy earned from below — is the difference between a deployment that builds calibrated trust and one that installs an impressive product and hopes for the best.

Applied to a McDonald’s franchise rollout, this architecture looks like four phases, each with a specific interface design.

Phase 0 — Shadow mode. The AI observes the building, records patterns, and builds its model. No autonomous actions. The interface in this phase makes the learning visible: what seasonal patterns is the system identifying? Where does the Chicago location’s kitchen exhaust load profile differ from other sites in the network? Operators who watch a system learn alongside them develop a more accurate mental model of what it understands and where its blind spots are. Shadow mode is not a data collection phase that users tolerate before the real product starts. It is the beginning of calibration.

Phase 1 — Recommendation mode. The AI proposes; the operator decides. The interface design here carries a critical requirement: declining a recommendation must be as fast and frictionless as accepting it. No confirmation dialogues for declines. No documentation requirements. No warning that declining will “reduce system performance.” If override is harder than approval, automation bias has been designed into the workflow, and the approval clicks that follow are not meaningful oversight — they are the appearance of oversight.

Phase 2 — Supervised autonomy within explicit boundaries. The AI acts automatically, but only within a setpoint range that the operator has reviewed and pre-approved. The interface shows the boundary itself — not just whether the AI is currently inside it, but what the boundary is and why it exists. As performance is demonstrated within those limits, the operator can expand them deliberately, with an understanding of what is being agreed to. In the Merck deployment, Phaidra began with a one-month recommendation-only demo before any autonomous control was granted. The results from that demo — 16% energy savings, 70% improvement in thermal stability — were what justified expanding the scope. The data made the case; the operator made the decision.

Phase 3 — Full autonomy, earned. At this stage, the interface shifts its primary function from oversight of individual decisions to exception management and retrospective review. The AI decision feed — a plain-language, timestamped log of every action taken, its stated reasoning, and the input data that drove it — becomes the primary accountability artefact. The operator’s daily relationship with the system is reviewing what happened yesterday and flagging anything that doesn’t make sense, not watching a live dashboard hoping nothing goes wrong.

The compressor fault in Chicago cost the franchise somewhere between $8,000 and $25,000 in emergency repair and lost morning service. Across 1,400 locations, a 2% annual incident rate at that cost level is a material liability. The interfaces that would have prevented it — sensor health visibility, unambiguous override scope, and an action log readable by an operator — are not expensive to build. They require product decisions made earlier in development, when the team is scoping what belongs on the main screen versus what belongs in the maintenance portal.

Phase 0 to Phase 3 (4 Steps). The path is the product.

Phase 0 to Phase 3 (4 Steps). The path is the product.

The patterns that separate trustworthy systems from systems that look trustworthy

After examining a range of building AI deployments — successful ones and instructive failures — a consistent gap emerges between products that build calibrated trust and products that perform well until they don’t.

Products that build calibrated trust share several design commitments. They display AI confidence at the decision level, not the system level — not “94% accurate overall” but “this recommendation: low confidence, input sensor outside normal range.” They surface input health — sensor status, data freshness, model confidence — on the same screen as the outputs those inputs are driving. They make override immediate and frictionless, and treat override history as a useful signal rather than a user error. They maintain a plain-language action log that any operator can read and question without engineering support. And they name their operating modes with precision, so “manual” means fully manual and “pause” means the entire system is paused.

Products that fail while appearing to succeed share different commitments. They aggregate uncertainty into system-wide accuracy numbers that obscure per-decision confidence. They make override feel consequential — confirmation dialogs, performance warnings, documentation requirements — training operators to avoid it until the moment it genuinely matters. They flood operators with undifferentiated alerts carrying no root-cause context, producing the well-documented “cry wolf” effect in which the one alert that matters eventually gets processed the same way as the 300 that didn’t. And they name themselves as more autonomous than they are, creating exactly the miscalibrated trust that makes the 5% scenario so costly.

The Tesla Autopilot case is the most public version of this anti-pattern at scale. In December 2025, a California administrative law judge found that the name “Autopilot” followed what was described as a long tradition of using ambiguity to create inappropriate trust. Tesla subsequently dropped the standalone Autopilot branding. The building automation industry faces no equivalent regulatory pressure on naming yet — but the structural problem is identical: products marketed as “fully autonomous” operating in modes that require active operator attention, borrowing trust they have not earned.

Compliance as a design brief, not a legal requirement

The EU AI Act’s requirements for high-risk AI systems — applicable from August 2026 — are worth reading as a design specification rather than a compliance checklist.

Article 13 requires that high-risk AI systems be designed so their operation is “sufficiently transparent to enable deployers to interpret a system’s output and use it appropriately.” The Chicago interface failed this standard. The franchise manager could not interpret what the system was doing at 7:00 AM with the information her screen provided.

Article 14 requires that operators be enabled to “remain aware of automation bias, correctly interpret outputs, decide to disregard or override outputs, and intervene.” The override that wasn’t fully an override failed this standard. The manager believed she had intervened. The system continued acting.

The regulation explicitly names AI in the management and operation of critical infrastructure — including heating and electricity — as a high-risk category. Whether any specific building AI product falls inside that boundary depends on implementation details still being clarified in enforcement practice. But the design implications are clear regardless of where the legal line lands: systems need to make their limitations visible, make override genuinely effective, and treat operator understanding as a design goal rather than a training problem.

The NIST AI Risk Management Framework’s seven trustworthiness characteristics offer a useful design audit for any team building in this space. The characteristic most consistently underdeveloped in commercial building AI is “explainable and interpretable” — the distinction between a system whose mechanism can be described in documentation and a system whose outputs mean something actionable to a specific operator at a specific moment. That distinction is what separated the platform that performed well in pilot testing from the one that produced two compressors in fault state before the breakfast rush was over.

What the Chicago failure actually teaches

The scenario that opened this piece ends with a repair bill, a lost morning of service, and a franchise owner who is considerably more skeptical of the AI platform than she was the day before. In isolation, it’s a bad day. Across a network of 1,400 locations, it’s a product design problem with measurable financial consequences.

But the deeper lesson isn’t about cost. It’s about what “trustworthy” means when it’s applied to an autonomous system operating in a complex physical environment with non-expert operators.

Trustworthy doesn’t mean accurate in ideal conditions. Trustworthy means the system provides what operators need to make good decisions in non-ideal conditions, which is the only time those decisions actually matter. A platform that performs beautifully in stable conditions and becomes illegible when a sensor drifts is not trustworthy. It is reliable until it isn’t, with no mechanism for operators to know which state they’re in.

The design goal worth pursuing, and worth evaluating products against, is not confidence. It’s calibration. The ability for an operator, at any moment, to have an accurate sense of what the system knows, what it doesn’t know, and what she should be watching. That’s not achieved through better algorithms. It’s achieved through better decisions about what belongs on the main screen, what “manual” means in an override interface, and what an operator actually needs to hear at 7:00 AM when something is starting to go wrong.

The Chicago franchise manager did nothing wrong. She responded rationally to the information her interface provided. The information the interface provided was insufficient. That’s a product problem. It has a product solution.

The goal was never maximum autonomy. It was appropriate autonomy and calibrated to what the system has earned and what the operator can actually see.

The goal was never maximum autonomy. It was appropriate autonomy and calibrated to what the system has earned and what the operator can actually see.


메타데이터
post_id
3eaeddf8f54a
slug
when-the-ai-was-right-and-the-optimization-still-went-wrong-3eaeddf8f54a
url
https://medium.com/@hamedsattarian/when-the-ai-was-right-and-the-optimization-still-went-wrong-3eaeddf8f54a
canonical_url
https://medium.com/@hamedsattarian/when-the-ai-was-right-and-the-optimization-still-went-wrong-3eaeddf8f54a
author_url
https://medium.com/@hamedsattarian
status
ok
fetched_at
2026-06-26 03:39:16