← Back to list

The inevitability of AI failures.

When predictable, preventable AI mistakes end up in court.

UpstreamThoughts · 2026-08-03 08:41 · 1 claps · 14.8 min read
#ai-failures #designingai #ai-product-management #artificial-intelligence #ux-design
Open on Medium ↗
Wiki topics: AI · AI · General UX · UI/UX Design BIZ · Business Strategy 📋 · Product Management ⚖️ · Law & Justice

The inevitability of AI failures.

When predictable, preventable AI mistakes end up in court.

Spotless. That is how Tripadvisor’s AI summarises the cleanliness of the Riu Palace Santa Maria in Cape Verde. The same page holds 32 one and two-star reviews posted between December 2025 and April 2026. 14 describe at least 1 member of the party falling seriously ill with food poisoning. 1 guest died. At least 412 holidaymakers who say they fell ill after staying there have since joined a group legal action, with 7 deaths reported since 2023. This prominent summary in question sits directly above the reviews it was built from.

Which? published this in an investigation this month, and it gets worse the further you read. Guests reported raw chicken, birds in the buffet, dead mice by the seating area. Tripadvisor’s trip-planning bot, Ollie, asked directly about the risk of food poisoning at the resort, calling it “quite unlikely”. At a hotel in Turkey where several reviewers described repeated sexual harassment by staff, including two accounts of a staff member following guests’ daughters, the AI summarised the service as “friendly”.

… reviewers described repeated sexual harassment by staff, the AI summarised the service as “friendly”.

You can probably picture the conversations that led to shipping this — an AI summary feels like the safest possible first AI feature: it touches no transactions, sits on top of content you already have, and the reader can cross-reference it. Worst case, users scroll past it — what’s the harm? Well, to start with, that sort of reasoning misses how a summary is perceived. It is a presumed value-add, right? It’s doing a job of saving time for the reader by highlighting the most relevant information. It is also a claim made in your product’s voice, prominently placed at the top of the page not necessarily because the decision makers trusted its quality, but because someone somewhere needed to report on page views, and top of the page guarantees eyeballs on this hot-off-the-press cutting-edge feature. But to the same eyeballs that very same summary carries the brand’s credibility, after all according to the brand itself “Tripadvisor [is] trusted by millions of travelers for over 25 years”. So it is not unreasonable to assume that to a user faced with the option to manually sifting through reviews, they would default to trusting the summary. To a user who trusts the brand it is the feature helping them decide.

This failure, and many like it, was not inevitable. It’s not something you can brush off with “there was no way to test this”. Which? checked Google’s AI overview of the same resort. It warned of “potential for illness” and flagged outbreaks. Two systems attempting to assist with the same task, similar raw material, and only one treated severe signals as something to surface rather than to average out. The difference between the two features is a set of product and design decisions these companies made, explicitly or implicitly — how and whether to weigh severity, when and how much to test and whether to test at all. One team spent the time on quality and prioritised user safety while the other didn’t.

Tripadvisor’s response, when challenged, did not help their case. The company said it fundamentally disagreed with the investigation, that the summaries were doing what they were designed to do, and that users could check them against the reviews. The last point being the weakest of the three, like leaving a kid in a candy store and assuming they could just choose not to eat it. And to me it pokes at a bigger question — who bears the responsibility for AI slop, and should it ever be the user if the user cannot opt out of the unreliable, even potentially harmful features?

Later in their article Which? pointed out the obvious — Tripadvisor chose to put the summaries at the top of the page, and Which? called the failure to surface critical safety information potentially life-threatening. The travel blogger at Head for Points, who is what you might call a professional traveller and had liked the feature at the start, presumably for its promise of saving time, wrote: “I thought it was genuinely useful. I should have known better.” And I would hazard a guess after reading this, you’re probably questioning Tripadvisor’s trustworthiness even though you’re likely not affected. Call me hyperbolic, but reading the Which? article alone is enough for me to not want to trust Tripadvisor again. They have positioned themselves as an authority in the travel space, it’s in the name — Tripadvisor; a brand with quarter of a century experience, on a mission “to be the world’s most trusted source for travel and experiences”. It is not unreasonable for their traveller users to have an expectation of reliable, trustworthy information. Now, I don’t work there and don’t know the background to how this feature came to be, so I can only speculate. But such a poorly calibrated AI summary feature reads far less as a feature intended to help travellers save time and find trustworthy information, and far, far more as an attempt to suppress negative feedback in an attempt to increase the conversion rates for the businesses that list with them, and in turn increase their own profit. Alas my personal trust tether is fairly short, so fixing the summaries will not bring the brand back for me.

The asymmetry of trust

None of this should surprise anyone who has read the trust literature. It’s a well studied, well documented, knowable concept that — at least in theory — Product decision makers at companies like Tripadvisor have at their disposal to evaluate.

Trust in general, and trust in a non deterministic AI feature specifically, accumulates very slowly through hundreds of unremarkable correct answers and uneventfully boring yet reliable interactions. The same trust if taken for granted can evaporate in one sloppy half-baked mistake. And because these failures follow patterns researchers have documented for over thirty years, they are predictable. Paul Slovic described the asymmetry principle in 1993, long before LLMs. Negative events damage trust far more than positive events build it. Trust forms slowly, can be destroyed by a single incident, and once destroyed may never fully return.

Trust accumulates very slowly through hundreds of unremarkable correct answers and if taken for granted can evaporate in one mistake.

Berkeley Dietvorst and colleagues ran a more recent experiment in 2015. People who watched an algorithm and a human make the same mistake lost confidence in the algorithm faster, and then refused to rely on it even after seeing it outperform the human. We forgive people their errors but we do not extend the same grace to machines.

Disuse — a word coined by Human factors researchers Parasuraman and Riley back in 1997: the abandonment of automation by the people it has burned. And mark my word, I think it will have a resurgence.

Put those together and the Head for Points arc starts looking less like an anecdote and more as a documented trajectory that was predictable. Predictable means preventable, and preventable means shipping features like this anyway is a choice. A choice — whether explicit or implicit — to deprioritise the safety and wellbeing of their traveling users and ultimately their trust in exchange for… what? Saving a week of researcher and developers’ time not testing and refining it? Or happier businesses that advertise with Tripadvisor? We will probably never know, but in any case, it is a choice, not a surprising, unpredictable inevitability.

The last two years in production

If the lab evidence feels distant, the field evidence is now arriving on a schedule. The excuses they’ll come up with just to avoid doing the hard work are tiring: “Vibe coding and founder mode means we don’t need engineers!”, “It’s cheaper to build and ship than it is to research. Discovery is dead!”, “We can’t possibly predict all the ways they’ll use this, testing is dead too. Test in prod!” I think I’ve covered most of the trending nonsense.

Air Canada’s website chatbot invented a bereavement refund policy that never existed. A grieving customer relied on it, was refused, and took the airline to a tribunal. Air Canada argued, in effect, that the chatbot was (and I quote) “a separate legal entity that is responsible for its own actions.” The tribunal member luckily saw through the BS and called that “a remarkable submission” and made the airline pay. The fare difference was about C$650, immaterial to the P&L, but the lost case and its defence became a global case study in how not to own your product’s words. Shameful abdication of responsibilities.

Cursor, an AI coding company, is a great example of AI confabulations as a business risk and gives us the cleanest measurement of how fast disuse happens. In April 2025 its AI support bot invented a policy limiting subscriptions to one device, delivered with corporate confidence as a “core security feature”. Developers cancelled within hours. One wrote that his workplace was removing the product entirely. Cursor’s co-founder apologised, refunded the user, and started labelling AI-generated replies. And while damage control and the response was decent and came relatively fast, it arrived after the exits. Switching costs were near zero, so the trust break led to high churn in a matter of hours.

Apple shipped AI notification summaries, and by January 2025 was pulling them for news after the feature fabricated headlines, including a false claim about a murder suspect that appeared under the BBC’s name. One of the most trusted consumer brands in the world could not safely summarise news on a lock screen, and every wrong headline carried both Apple’s authority and the BBC’s.

Examples are a-plenty. Different logos, different companies, different customers, same story. Most of these arrived after the earlier ones had already made global headlines. Predictable and preventable, and yet teams are told to “ship it” anyway.

Who’s culpable then?

Earlier I asked who bears the responsibility for AI slop. At the moment neither side of the exchange — users on the one, companies and their AI feature building teams on the other — is doing its share, and the two sides are not equally placed to fix it.

Take the user side first. While companies race to ship subpar AI at speed, it can be comforting to believe the people who consume AI outputs compensate with their skepticism and technical literacy, making it impossible to truly mislead them — the unpalatable justification made by Tripadvisor. And yes, the AI revolution is here, the genie is out and we aren’t about to time travel back to simpler times of pre-2020 unsexy AI world. So the general population has a lot of technical gaps to fill and upskill to do in order to coexist more safely with this technology. However, to expect a user to anticipate where the proverbial skeletons are buried is entirely irresponsible and frankly the tech industry should be held to a higher standard.

66% [of AI users in work setting] rely on AI output without evaluating its accuracy

The University of Melbourne and KPMG surveyed over 48,000 people across 47 countries and found that 70 percent of people believe AI regulation is required and majority don’t believe that what we have now is sufficient to make AI use safe and protect people from harm, and only 46% are willing to trust the technology. That study has more alarming findings — at work the picture is worse. 66% rely on AI output without evaluating its accuracy, 56% have made mistakes because of it (that they know of. The figure is likely higher in reality), and the study team’s own writeup adds that most employees hide their AI use, with over half presenting AI content as their own. And what the overwhelming majority said they need in order to be more willing to trust the AI system is for companies to do more to ensure outputs are reliable, with independent third parties assuring it isn’t just fluff. To me that reads as more than just a label or disclaimer doing the heavy lifting of accountability.

Gillespie, N., Lockey, S., Ward, T., Macdade, A., & Hassed, G. (2025). Trust, attitudes and use of artificial intelligence: A global study 2025. The University of Melbourne and KPMG. DOI 10.26188/28822919

Microsoft and Carnegie Mellon researchers documented another mechanism: across 319 knowledge workers, the more confidence a person had in the AI, the less critical thinking they applied to its output. This is the origin story of your internal AI slop cannon. The verification layer everyone assumes exists is mostly missing. The builder skips it to ship faster and the user skips it because the output looks finished so the automation bias carries them the rest of the way. Add well-documented tokenmaxxingmade worse by Goodhart’s law, all of it likely masquerading as “AI productivity gains”. It is all a big mess.

The corporate answer, increasingly, is a label. By now you’re probably so familiar with it, you’re blind to it, but it’s the tag that something is “AI-generated”, add “may contain mistakes”, maybe a link to the underlying sources, and the risk feels handled. Tripadvisor’s response to Which? reads as almost farcical through the culpability lens: its community, the company said, “has the common sense to check any AI advice” against the reviews underneath. Cursor’s fix, after its support bot invented a policy, was to label AI replies. In Europe, as of 2 August, the label will stop being a viable choice. The EU AI Act’s transparency rules will require AI systems that interact with people to say so and synthetic content to be marked, duties that reach global businesses serving EU users.

I’m not naïve. Companies have not convinced themselves the label is a customer-first approach to this whole mess. The label is convenient. It’s easy. It feels like doing something about the problem caused by not doing more. But the label is only addressing the symptom, it’s not solving anything.In fact, evidence shows that the label does not lead to more people verifying labeled outputs. The 66% from before, who are skipping evaluation are already using tools that carry disclaimers. It did not shield Air Canada, whose tribunal held the company responsible for everything on its website and refused to accept that customers “should have to cross-check one part of a site against another”. That is as close as a tribunal has come to answering the opt-out question: due diligence does not transfer to the user just because the company points at an AI label. Nor does a label protect trust itself. Schilke and Reimann ran thirteen preregistered experiments and found that disclosing AI use lowers trust in the discloser, whether the disclosure is voluntary or mandated. The label is at most a CYA activity, never a shield when the output is wrong. The same research found not disclosing AI is worse once it’s discovered as such, so concealment is not the answer either. The answer is the unfashionable first-principles work: make the damn thing good enough that the label reads as an informative description instead of a warning. Oh and, in case it’s not obvious — I mean good enough for the user, not just the PowerPoint decks.

Because the outcome, if nothing changes, is already measurable. Pew’s February surveyfound that among users who have not adopted chatbots, 76% say they do not trust the tools to give accurate information. What about the other end of the spectrum? Surely the ‘techies’ get it? I found this equal parts interesting and infuriating gem in The Marketing Journal. Let’s pause on the title: “Lower Artificial Intelligence Literacy Predicts Greater AI Receptivity”. The work shows that higher AI literacy is associated with lower AI adoption. The researchers conclude that “efforts to demystify AI may inadvertently reduce its appeal, indicating that maintaining an aura of magic around AI could be beneficial for adoption.” So the more you understand the technology, the less magical it looks against the marketing, so tech savvy users see through the flashy promise of technological utopia and do not adopt the tech either, leading to various degrees of automation aversion. (Side tangent to the reader, that marketing journal also tells me based on their intent and their conclusions about maintaining the “magical aura” that companies are catching on to it and are investigating ways to increase marketing efficiency and public perception of AI, even Sam Altman has spoken of AI having a Marketing and PR problem).

The more you understand the technology, the less magical it looks against the marketing.

We have covered a lot of ground, so a quick recap. Market pressure leads companies to ship fast and be seen to adopt cutting-edge technology. To limit the blast radius, they attach labels and disclaimers meant to shift the cognitive effort of output reliability onto the user. Users who haven’t adopted the tech avoid it because a) they doubt the outputs and/or b) they understand constraints and aren’t compelled by the marketing. That all reads to me like free fall towards disuse, caused by rushed, poorly scoped AI slop shipped despite the abundance of available lessons learnt by the industry thus far. A company that thinks quality can be substituted with a label instead of testing has written down that it knew the output might be wrong, and shipped it to the top of the page anyway. Predictable and preventable.

What’s the hidden cost of speed?

The trade every scoping conversation is making is value against risk, and the two sides get modelled unevenly. Value can be expressed as value to the user or the business. The latter is easier to weaponise — it’s the wasted engineering hours, or the extra testing efforts, or the “time we can spend on shipping the other features”. Risk on the other hand is harder to put a KPI on if it isn’t something you’re used to doing. So it’s much harder to advocate for, especially if it’s a risk to user trust. Afterall — how can you put a number on the thing you prevented from happening to people who aren’t convinced it was going to happen? This is likely all too familiar if you’re in a UX role. But no matter how difficult it is to articulate or accurately predict, it’s still worth pumping the brakes to ask the questions about what are we assuming. If we believe it is safe to ship a half baked AI feature because all users will meticulously read reviews to verify the summaries, the cost of evaluating that assumption is as cheap as a single research study. So surely that can’t have been the reason Tripadvisor chose to launch the feature as is. So why did Tripadvisor gamble? Whatever it is, my guess is they thought that a shiny roadmap line item and lower short-term cost was the quick win. And time will tell, but the quick win might just have taken for granted the trust it took twenty-five years to earn.

So… what now?

Where do we go from here? Are trust debt and the cost of speed two sides of the same coin? You tell me. We are undeniably in the thick of it now, everyone is seemingly doubling down on AI-everything and the opportunities are really tempting. The tech of course can bring a lot of value. The LinkedInfluencer tech bros posting about fleets of agents running their one-person business while they sleep increases pressures on decision makers to AI harder, faster, better, stronger. The market pressures force decision makers to prioritise in ways they may otherwise not have, balancing shareholders, time-to-market rat race, PR and so forth. We are all in the hot seat.

But what is equally as undeniable is the risk that these opportunities introduce. The sloppy, rushed work brings with it increased, highly visible and potentially devastating consequences, like for some of Tripadvisor users.

So what do we do? Governments are paying attention and new policies are in the works, which no doubt will help to pump the brakes. But law moves slowly and is reactive, often adjusting after it’s too late for many. In the meantime, I’d summarise it by repeating what I said before. Make the damn thing good enough for the user. Which means test the damn thing with the user — not just against your internal opinions — before you ship the feature. That sure beats the popular yet miscalculated ‘eFFF around find out’ approach every time in the long run.

Make the damn thing good enough for the user. Which means test the damn thing with the user

I’ll leave a short list of a few threads you to pull on to explore ways you could help advocate for the predictable to be preventable.

  • A surprisingly simple way to make the choice explicit, bring out the big guns, the iron triangle to help articulate the tradeoffs. Here’s a simple yet playful tool that’s a good way to structure conversations about what the industry obsessed with new shiny object is optimizing for and what it might cost us.
  • Another time tested option for you to explore is Operational Design Domain, if you’re a system thinker, you will find this to be comforting, familiar and intuitive.
  • Speaking of Systems Thinking, it too offers many techniques and approaches that make the invisible visible, and, dare I say it one last time, the predictable preventable.
  • If you know me, you know I’m always dying on some hill about the importance of defining why you’re building what you’re building, and defining success metrics and the spec of what you’re building before you build it, a.k.a build it on paper before you build it, you’d be surprised how much further and faster you’ll get. There’s art to the craft and to me no one does it better than John Cutler. It’s hard to point to any one piece, but start here and choose how deep your rabbit hole goes.

You will be unsurprised, I hope, by how un-hypeworthy these things are. There’s no need to reinvent the wheel every time new tech comes along. I have found in practice that sticking to tech-agnostic, human-centred and time tested concepts will get you most of the way there, and testing with real users will carry you over the finish line.

Until next time.

~

If you enjoyed this, follow me on Substack, which is where I post all my full length posts first

All opinions and views are my own and do not belong to an LLM or reflect the views of the institutions(s) I choose to work for.


메타데이터
post_id
99ff048edd58
slug
the-inevitability-of-ai-failures-99ff048edd58
url
https://medium.com/@upstreamthoughts/the-inevitability-of-ai-failures-99ff048edd58
canonical_url
https://medium.com/@upstreamthoughts/the-inevitability-of-ai-failures-99ff048edd58
author_url
https://medium.com/@upstreamthoughts
status
ok
fetched_at
2026-08-27 13:14:59