← Back to list

Generative AI Has a Patient Safety Problem.

AI is in clinical care. Patient safety frameworks are not.

Esther Olowoloba in The Systems Rewrite. · 2026-05-06 13:00 · 1 claps · 17.1 min read
#generative-ai-use-cases #patient-safety #ai-in-healthcare #health-equity #health-technology
Open on Medium ↗
Wiki topics: SAF · Safety & Alignment AI · AI · General DH · Digital Health & Health Tech ✊ · Equality & Identity

Generative AI Has a Patient Safety Problem.

AI is in clinical care. Patient safety frameworks are not.

I would like this article to start from a personal place.

I trained and worked as a registered nurse in Nigeria, and if you have not worked in a Nigerian public health facility, let me paint a picture for you.

You are managing more patients than the ward was designed for. The monitor alarms are going off for three beds simultaneously. The documentation is paper-based or sometimes electronic, but with limited technology, the pharmacy is two floors down, and the relatives of every patient in the room are watching you with the specific anxiety of people who know that the system is under-resourced and who are counting on you personally to be the thing that does not fail.

In that environment, patient safety is not an abstract framework. It is the calculation you are making in real time, constantly, with incomplete information and too many competing demands.

I began a career in digital health policy because I believe technology can change these problems. I believe AI, used responsibly, can reduce the cognitive load on clinicians, catch things human eyes miss, and extend the reach of healthcare to people who currently have no meaningful access to it, and I genuinely believe this.

But I have also watched the conversation about AI in healthcare become increasingly detached from the realities of clinical environments, specifically those in Africa. We talk about what AI can do, its efficiency, early detection, and reduced clinical burden. We do not talk enough about what happens when governance structures are absent, when AI fails, and when the person bearing the consequences of that failure is the most vulnerable patient in the most under-resourced setting.

That is what this article is about. Not whether generative artificial intelligence (AI) belongs in healthcare, it does, but what patient safety actually means in the age of generative AI, and whether the frameworks we are building are adequate for the environments where the stakes are highest.

This is the sixth article in my policy prescription series. In the first, I argued that the interoperability problem is political, not technical. In the second, I showed that health systems create the conditions for misinformation to thrive. In the third, I said health data does not belong to patients because the policy built it that way. In the fourth, I showed that AI bias is a health equity crisis. And in the fifth, I argued that healthcare needs to borrow financial services’ governance model before it creates its own version of 2008.

Each of those articles was building toward this one because they all described the conditions that make patient safety failures in the generative AI era not just possible but predictable.

Let’s begin with some facts

The Emergency Care Research Institute (ECRI), one of the most respected patient safety research organizations in the world, named AI chatbot misuse the number one health technology hazard for 2026—the number one technology hazard in healthcare this year.

To understand why that ranking matters, you need to understand what ECRI is actually saying. It is not saying AI chatbots are inherently dangerous. It is saying they are being deployed at a scale and in clinical contexts that their safety profiles do not support, without the governance infrastructure to catch failures, and with a user population that largely does not understand what these tools can and cannot do. Every day, more than 40 million people turn to ChatGPT alone for health-related answers. And ECRI’s description of why this is dangerous is worth quoting in spirit, if not in exact words. These systems are programmed to sound confident and always to provide an answer, even when the answer is unreliable. Medicine is a fundamentally human endeavor, and the algorithm cannot replace the expertise that the clinician brings.

Now here is where please stay with me, because this is where the conversation usually stops, but where I think it actually needs to start.

Those 40 million daily ChatGPT health queries are global figures, and they include millions from people in Nigeria, Ghana, Kenya, South Africa, and every African country where smartphone penetration is rising, and access to a qualified clinician is not.

For those users specifically, the chatbot is not a convenience. It is often the only accessible source of health information available to them. And for those users specifically, the consequences of a chatbot that confidently provides wrong triage guidance, fabricates a drug interaction, or generates a plausible-sounding but dangerous treatment suggestion are not inconveniences and can be fatal.

A 2026 BMJ Open audit found that 49.6% of AI chatbot health answers were problematic. That is an extraordinary failure rate for any clinical tool, and it would not be acceptable in any other context. It is being tolerated in this one because the tools arrived faster than the policy frameworks designed to govern them, and because nobody has yet created the accountability structures that would make tolerating it consequential for the institutions deploying these tools.

Patient safety, what exactly does this mean?

We constantly say “patient safety” in healthcare. It is on hospital walls, in accreditation frameworks, in mission statements, in the job descriptions of entire departments. But it has a specific technical meaning, and I want to be precise about it because generative AI is testing that meaning in ways the system has not yet fully acknowledged.

Patient safety means protecting people from the gap between what a healthcare system intends to do and what it actually does. It is the set of practices, systems, and accountabilities designed to catch errors before they reach the patient, to identify them when they do, and to prevent them from recurring. It is built on a foundational assumption that errors have traceable causes, identifiable accountabilities, and learnable lessons.

Generative AI challenges every one of those assumptions, and I want to explain how, specifically, because this is not an abstract concern. It is the practical reality of what happens when large language model (LLM)-based tools enter clinical settings without adequate governance.

The first challenge is the scale of error. A nurse who gives a patient incorrect information about their medication is making one error for one patient in one interaction. A patient-facing AI chatbot that systematically under-triages emergencies is making the same class of error across millions of interactions, with no escalation mechanism and no feedback loop. A Mount Sinai study published in February 2026 found that ChatGPT Health under-triaged 52% of genuine emergencies, often steering people away from urgent care when they needed it most. That is not a documentation problem but a life-or-death decision made incorrectly at a mass scale, and the patient safety frameworks built around individual clinical errors are not designed to catch it.

The second challenge is confident misinformation. This is what the AI industry calls hallucination, and I think that word is too gentle for what it describes in a clinical context.

A hallucination is not a random error. It is a confident, coherent, grammatically correct, clinically plausible piece of information that is factually wrong. An AI that hallucinates a medication allergy, a past diagnosis, or a drug dosage does not flag itself as uncertain and delivers the fabricated information in the same register as accurate information. Clinicians and patients have no internal signal to distinguish the two.

A 2025 study found that 43 million patients in the United States alone ask chatbots medical questions at least once per month, and that unsafe advice has been documented across multiple clinical domains, including emergency triage, medication guidance, and diagnosis. The study found ChatGPT had misdiagnosed shingles as ringworm and given false reassurance to a patient suffering a transient ischaemic attack, delaying their care.

The third challenge is diffuse accountability. When an AI documentation tool generates an inaccurate clinical note and a clinician acts on that note, and a patient is harmed, who is accountable? The clinician who missed the error? The hospital that deployed the tool? The vendor that built the model? In traditional patient safety, accountability is traceable. In AI-mediated care, it is diffuse by design, and that diffusion is a governance consequence rather than a technological inevitability. It persists because no jurisdiction has yet established clear rules about who owns the outcomes of AI-generated clinical content. In the meantime, patients bear the consequence, and nobody bears the accountability.

The fourth challenge is the erosion of clinical judgment as a safety net. Clinical judgment is patient safety’s last line of defense. When every other system fails, the experienced clinician at the bedside catches it. A survey by National Nurses United, the largest association of registered nurses in the United States, found that AI technology often contradicts and undermines nurses’ clinical judgment and threatens patient safety. If the system is designed to consistently override clinical judgment in favor of AI output, the safety mechanism that catches AI errors is exactly the mechanism the AI is replacing. That is a structural vulnerability that should alarm anyone who understands how patient safety systems actually work.

Where generative AI sits in clinical care right now

Generative AI in healthcare is not one thing. Understanding the patient safety implications requires identifying the specific clinical context you are talking about, because the risks differ and the policy responses need to be different, too. The Institute for Healthcare Improvement (IHI) Lucian Leape Institute convened an expert panel in January 2024 to examine the risks posed by generative AI for patient safety. It identified three primary use cases: documentation support, clinical decision support, and patient-facing chatbots. Let me take each one seriously.

Documentation support is where the latest deployment is occurring. Ambient AI tools listen to clinical encounters and automatically generate clinical notes, discharge summaries, and patient communications. Nurses spend between 19% and 35% of their working hours on documentation, and tools that reduce that burden could meaningfully increase time for direct patient care. I understand the appeal intimately. I have spent hours documenting in conditions where the documentation itself is a safety risk, because it pulls your attention away from the patient. So I am not dismissing this use case.

But I need to say clearly what the risk is: an AI that automatically completes note-taking can miss important clinical nuances, generate inaccurate summaries, or introduce details the clinician never said. And a wrong note is not just an administrative problem. It is the information the next clinician will use to make decisions about that patient, possibly at 3 am when they have no context and are relying on the record.

Clinical decision support is where the stakes are highest. AI tools that synthesize patient data and produce diagnostic or treatment recommendations are operating in a territory where errors have immediate clinical consequences. The IHI panel described these tools as capable of saving clinicians time and improving diagnostic accuracy, while also noting they have flaws that may compromise patient safety. Both things are simultaneously true, which is exactly what makes this space so hard to govern. The tool can be right or wrong in ways that appear identical to the clinician relying on it.

Patient-facing chatbots have the greatest exposure, and this is the use case I am most concerned about in the African context. When a patient in Lagos, Kano, or Ibadan uses an AI chatbot to ask about their symptoms, interpret a test result, or decide whether their chest pain requires a hospital visit, they are receiving health guidance from a system that has no access to their clinical record, no awareness of their medical history, no calibration for the disease profiles of their specific population, and no regulatory obligation to be accurate.

The chatbot sounds confident because it is designed to sound confident. Confidence is a function of how large language models are built, not a reflection of clinical certainty. And in an environment where the alternative to the chatbot might be no health guidance at all, the patient has no basis to discount its authority.

What This Means for Africa

I need to spend time here because the global conversation about generative AI and patient safety almost entirely assumes a high-income country context: sophisticated hospital systems, regulated clinical AI tools, and a patient population with meaningful access to alternative healthcare options. That assumption does not hold in Nigeria. It does not hold across most of sub-Saharan Africa. And the failure to disaggregate the conversation means we are building governance frameworks for the wrong patient.

Nigeria is the most populous country in Africa and one of the fastest-growing digital health markets on the continent. In May 2023, Nigeria launched NIGCOMHEALTH, described as the first African national digital healthcare platform, a telehealth service platform designed to expand access to clinical services. The Nigeria Digital Health Initiative (NDHI), launched by the Federal Ministry of Health and Social Welfare, aims to create a national digital health backbone that integrates clinical, laboratory, and administrative data. Nigeria’s 2024 National AI Strategy emphasizes human-centered design and cultural sensitivity. These are real and meaningful policy commitments.

But a 2025 narrative review of AI adoption in Nigeria’s healthcare system found that infrastructure gaps, limited digital literacy, and weak regulatory frameworks remain the primary barriers to safe and widespread AI adoption. The SPEC-AI Nigeria trial, which tested AI clinical decision support across Nigerian health facilities between 2022 and 2024, documented significant contextual challenges, including supply chain problems, inconsistent power and connectivity, and the reality that most AI tools being piloted in Nigeria were developed outside the country for patient populations that bear little resemblance to Nigerian patients.

These tools arrive with their biases already embedded, trained on datasets where Nigerians are largely absent, and deployed in a regulatory environment that has not yet built the infrastructure to evaluate their performance on local populations.

Across the continent, the picture is structurally similar. Africa CDC unveiled the Continental Health Data Governance Framework in July 2025, a unified framework for health data management across African member states, scheduled for endorsement at the African Union Summit in February 2026. The African Union’s Continental Artificial Intelligence Strategy provides a strategic blueprint for AI governance on the continent. More than half of African countries now have data protection laws in force. These are genuinely important steps, and I want to acknowledge them as such.

What they are not is a patient safety framework specifically for generative AI. They address data governance. They do not address what happens when an AI chatbot undertriages a maternal emergency in a community where the nearest obstetric unit is two hours away. They do not address what happens when an ambient documentation tool generates an inaccurate clinical note in a facility with one doctor and fourteen nurses covering a ward of 60 patients. They do not address accountability when a patient in Nasarawa State acts on AI-generated health guidance that contradicts the community health worker’s advice and suffers a preventable harm.

The frameworks that exist address data. The patient safety risk in the generative AI era is not primarily a data problem. It is a deployment, accountability, and clinical governance problem.

The African Union’s Health Strategy 2016 to 2030 commits African governments to equitable healthcare for all citizens by 2030 and aligns explicitly with Universal Health Coverage (UHC) targets. The WHO’s Global Strategy on Digital Health 2020 to 2025 emphasizes cross-country collaboration to ensure equity and access in AI-enabled care. Both are frameworks that articulate the right values. Neither provides a mechanism for ensuring that a generative AI tool deployed in an African clinical setting is clinically safe for African patients. That gap is not a future concern but a present one, because the tools are already being deployed.

When Governance Fails: A Real Example

I want to show you what the absence of patient safety governance for generative AI looks like in an institutional setting, not in theory but in a documented, published government report.

The United States Department of Veterans Affairs (VA) is one of the largest integrated healthcare systems in the world. It serves millions of veterans across hundreds of facilities. It has a sophisticated health informatics infrastructure and a dedicated research apparatus. It is, in many respects, better resourced to manage AI safely than most health systems globally. A January 2026 report from the VA’s own Office of the Inspector General (OIG) found that the Veterans Health Administration (VHA) does not have a formal mechanism to identify, track, or resolve risks associated with generative AI.

227 documented AI use cases are in active operation across the VHA, with no formal mechanism to identify, track, or resolve the safety risks of any of them. The OIG’s language is direct: the absence of this process precludes a feedback loop and a means to detect patterns that could improve the safety and quality of AI tools used in clinical settings.

Without a feedback loop, clinician concerns go unaddressed, near-misses go unreported, and patterns of error across facilities remain invisible. The tools continue operating, flagged by nobody, accountable to nothing, until a harm is serious enough to surface through a different channel entirely.

If this is what governance failure looks like at the Veterans Health Administration, I want you to think about what it looks like at a district hospital in Kogi State, or a primary health center in rural Zambia, or a community clinic in Accra that has adopted a telemedicine AI platform provided by a foreign technology company under a donor-funded program.

The institutional capacity to identify, track, and resolve generative AI safety risks does not spontaneously exist. It has to be built and done before the deployment, not after.

My Recommendations: What Needs to Change Now

I want to structure my recommendations differently in this article because I think the standard numbered policy list creates a false impression that these are separate problems with separate solutions. Instead, they are interconnected failures that require interconnected responses, and I want to explain how they connect.

  1. Clinical Risk Classification Is Non-Negotiable

Before any generative AI tool is deployed in a hospital, clinic, or patient-facing application, it must be classified by clinical risk level using a standardized framework. High-risk tools, specifically those that influence diagnosis, triage, medication, or treatment decisions, require pre-deployment clinical safety evaluation, mandatory human oversight, and post-deployment adverse event reporting. Lower-risk tools, specifically those handling administrative tasks with no direct clinical decision input, require lighter oversight.

The classification framework must be developed by clinical professional bodies, not by AI vendors. Allowing vendors to classify their own tools’ risk levels is a conflict of interest so obvious it should not need stating, and yet it is currently the default in most jurisdictions. The classification framework is the prerequisite for everything else, because you cannot govern what you have not defined.

2. Building the Reporting Infrastructure

You cannot fix what you do not report. Once you have defined what a high-risk clinical AI tool is, you can require that it has a formal adverse event reporting mechanism separate from the existing incident reporting infrastructure. Every healthcare organization operating clinical AI must maintain a dedicated reporting system for AI-related safety events, including near-miss incidents.

This is what the VHA Inspector General identified as missing. It is what the ECRI report implies is absent at scale. Without it, the 49.6% problematic AI health answers documented in the BMJ Open audit remain invisible to the institutions producing them. The data about what is going wrong exists at the patient level. The governance infrastructure to aggregate, analyze, and act on it does not exist.

3. Tell Patients What the Tool Cannot Do

Once you have a reporting infrastructure in place, you can establish mandatory disclosure standards for patient-facing AI health tools. Every patient-facing AI health application must carry three specific disclosures, shown before the first interaction:

  • The tool is not a licensed clinician, and its outputs are not medical advice.
  • That it cannot access the user’s clinical history and therefore cannot assess their specific situation.
  • That it should not be used for emergency triage, with explicit direction to emergency services.

These disclosures matter most in environments where the chatbot is not a supplement to clinical access but a substitute for it. In Nigeria and across Africa, where smartphone penetration is rising faster than healthcare infrastructure, tens of millions of people will make health decisions based on AI guidance before they consult a clinician. Those people deserve to understand what they are interacting with.

4. Clinical Voices Must Be Included in Deployment Decisions

The National Nurses United survey finding that AI tools contradict and undermine clinical judgment is not a workforce concern in isolation, but a patient safety concern. Nurses are the people who interact most directly with patients and with AI tools at the point of care.

They are also the people most systematically excluded from decisions about whether and how to deploy those tools. I have said this in this series before, and I will say it again: requiring nurses and frontline clinical staff to hold formal roles in AI deployment governance committees, with documented authority to flag safety concerns and pause deployment pending review, is not a soft inclusion measure. It is the operationalization of the clinical judgment that the safety literature identifies as essential.

5. Local Validation Before Local Deployment

For Africa specifically, the most urgent structural change is mandatory local validation as a condition of deployment for any AI tool imported from a high-income country context. The Continental Health Data Governance Framework, endorsed by the African Union, provides the data infrastructure foundation. What needs to be built on top of it is a clinical performance validation requirement: any AI tool that has not been tested on populations representative of the health facility’s actual patient population cannot be deployed in a patient-facing clinical context.

This is not a bureaucratic barrier to innovation. It is the minimum safety standard that a drug requires before it can be dispensed. An AI tool giving clinical guidance to patients in Abuja should be required to demonstrate that it performs safely for patients in Abuja, not just for patients in Boston or London, where it was trained and validated. The African Union’s Continental AI Strategy commits to this principle as an aspiration. It needs to be built into enforcement-level procurement requirements.

Underpinning all of this is the most foundational change: generative AI tools used in patient-facing healthcare contexts must be brought within the scope of clinical regulation in every jurisdiction. Currently, a chatbot that advises millions of people on whether their symptoms require emergency care operates in most countries with no clinical oversight requirements, no mandatory accuracy standards, and no adverse event reporting obligations.

The Food and Drug Administration (FDA)’s existing framework for software as a medical device provides a starting point. The EU (European Union) AI Act’s high-risk classification system provides another. Nigeria’s 2024 National AI Strategy and the National Health Insurance Authority Act 2022 provide the domestic policy architecture for anchoring national AI regulation. The specific decision on whether to bring patient-facing AI health tools within the scope of clinical regulation, with proportionate requirements that match the clinical risk level, is a matter for each jurisdiction. This is the foundational policy change, and everything else builds on it.

My Conclusion

I started this article by describing what patient safety feels like at the bedside of a Nigerian hospital ward. I want to end it there, too.

Because when I think about what patient safety means in the age of generative AI, I keep returning to that image: the nurse making real-time calculations, with incomplete information, too many patients, and the knowledge that if she gets it wrong, someone suffers for it. That nurse is the last line of defense, not because the system was intentionally designed that way, but because everything upstream of her has gaps that her clinical judgment is expected to fill.

Well-deployed generative AI should reduce the weight of those calculations. It should surface patterns she cannot see, document what she does not have time to write, and give her more of her attention back for the human being in front of her. That is the promise, and I believe in it.

But deployed without safety classification, adverse event reporting, disclosure standards, local validation, or clinical accountability structures, generative AI does not reduce the burden on that nurse. It adds to it. It introduces a new source of confident-sounding error into an environment that was already managing too many. And for the patient in the bed at the end of the ward in a facility that adopted an AI chatbot because it was free, came with donor funding, and looked like progress, it introduces a risk that nobody in that chain of decisions was required to account for.

Patient safety has always meant protecting people from the gap between what a system promises and what it delivers. Generative AI has not changed that definition. It has just made the gap much wider, much harder to see, and much more consequential for the populations who were already bearing the heaviest burden of the gaps that existed before it arrived.

If you are a patient anywhere in the world using AI tools for health information, please understand what they are and what they are not. They are language models trained to produce plausible text. They are not clinicians. If you are in any doubt about whether something is an emergency, contact a clinician or emergency services. A chatbot that sounds confident is not the same thing as a professional who has examined you.

If you are a clinician, your judgment is not a tie-breaker to be invoked when the algorithm is uncertain. It is the primary safety mechanism. Use it as such, and document it as such.

If you are building health policy in Nigeria or anywhere across Africa, the Continental Health Data Governance Framework is a foundation. What needs to be built on it now is a clinical AI safety layer: validation requirements, adverse event reporting infrastructure, and procurement standards that make local performance validation a condition of deployment rather than an afterthought.

If you are building or deploying generative AI in healthcare, the question is not whether your tool can help. Most of them are. The question is whether you have been required to demonstrate that it is safe for the specific population you are deploying to, and whether you are choosing to do so even when you have not been required to.

The answer to that question determines what patient safety actually means in the age of generative AI.

Thank you for reading. I will see you in my next article.

Bye for now.

Esther O.


메타데이터
post_id
d7ef0c744ddb
slug
generative-ai-has-a-patient-safety-problem-d7ef0c744ddb
url
https://medium.com/technology-with-starlife/generative-ai-has-a-patient-safety-problem-d7ef0c744ddb
canonical_url
https://medium.com/technology-with-starlife/generative-ai-has-a-patient-safety-problem-d7ef0c744ddb
author_url
https://medium.com/@estherolowoloba
status
ok
fetched_at
2026-06-09 15:37:30