← All news

Summit

Risks of AI in Healthcare: Ranked by Who Carries Them

EX Future Summit · 24 August 2026

The 2026 evidence on the risks of AI in healthcare, sorted by who absorbs the harm: patients, clinicians, institutions, and the risks nobody owns.

Risks of AI in Healthcare: Ranked by Who Carries Them

Almost every account of the risks of AI in healthcare is a list of categories. Bias, privacy, hallucination, opacity, job displacement. The list is accurate and it is nearly useless, because a category is not a risk. A risk is a harm of some size, landing on somebody specific, with someone else in a position to prevent it. Sorted that way, the same evidence produces a very different picture, and some of the items that dominate the disadvantages-of-AI listicles turn out to be small next to the ones nobody writes about.

So here is the 2026 evidence sorted by bearer. What the patient absorbs, what the clinician absorbs, what the institution absorbs, and the residue that has been assigned to no one at all. This assumes rather than repeats the adoption picture: if you want the deployment numbers first, start with how much AI is actually deployed in healthcare.

Risks the patient carries

Undertriage at the consumer front door

The sharpest measured danger in the current evidence is not in a hospital. It is in the tools patients reach first.

Researchers ran a structured stress test of ChatGPT Health using 60 clinician-authored vignettes across 21 clinical domains under 16 factorial conditions, producing 960 responses. Accuracy followed an inverted U. It peaked in the middle, at 93.0 percent for semi-urgent presentations, and collapsed at both extremes, to 35.2 percent for nonurgent cases and 48.4 percent for emergencies. Among true emergencies, 51.6 percent were undertriaged: patients with diabetic ketoacidosis or impending respiratory failure were directed to a 24 to 48 hour follow-up rather than an emergency department.

Two details make this worse than the headline. When a family member or friend minimised the symptoms, triage recommendations shifted, with an odds ratio of 11.7 in edge cases and the majority of shifts pointing towards less urgent care. And the crisis-intervention safeguards fired unpredictably across presentations of suicidal ideation, in one pattern appearing more often when no specific method was described than when one was. A guardrail that cannot be anticipated cannot be relied on, by the user or by anyone designing around it.

The asymmetry is the point. Overtriage costs an unnecessary appointment. Undertriage costs the window in which treatment works.

Recommendations that shift with the patient

Bias in this setting is not an abstraction about training data. It is a measurable difference in what the system recommends for two patients with the same presentation.

Evaluations collected by the US Agency for Healthcare Research and Quality found that leading models varied diagnostic and treatment recommendations by race, ethnicity, sex and socioeconomic status even on vignettes designed to minimise bias, generated racially stereotyped clinical scenarios, and reproduced refuted race-based medical claims about pain tolerance and kidney function. The mechanism is mundane and hard to remove: a model trained on historical hospital data learns the patterns in that data, including the ones that were never clinically correct, such as inferring sepsis risk from which ward a patient is in rather than from their physiology.

No route to contest the decision

The patient's third exposure is procedural, and it is the one the risk listicles never reach. A scoping review of 77 organisational AI governance frameworks found that contestability appeared in only 10 of them, 13.0 percent, with consent, confidentiality and medicolegal liability similarly uncommon. Only 10 frameworks carried all four components the reviewers considered necessary for real-world use, and the least common missing piece was an oversight mechanism. No framework in the review had been evaluated for whether it works.

Which means that for most patients, in most institutions, there is no defined way to ask why a model influenced their care and no defined route to redress if it did so wrongly. That gap belongs to the wider argument about the ethics of AI in healthcare, and it is not a philosophical one. It is an operational hole.

Risks the clinician carries

Error rates that assume a review step nobody has time for

Clinicians carry a different risk: they are the designated catcher of errors produced faster than they can be checked.

The rates are not marginal. Ambient digital scribes have been found to introduce errors in 70 percent of generated clinical notes. AI-drafted responses to patient messages carry a hallucination rate around 6 percent, with 7.1 percent of drafts judged to pose a severe risk of harm. A systematic review of more than 500 deep learning studies in radiology reported median diagnostic accuracy of 89.4 percent, though many of those studies lacked external validation or sat at high risk of bias.

Every one of those numbers is tolerable if a human reviews each output with full attention. None is tolerable if the tool was adopted precisely because there was not enough attention to go around. That is the trap, and it is structural rather than technical: hallucination is a consequence of models optimised to predict the next plausible token rather than to verify a fact, which is why confident phrasing and factual accuracy come apart so cleanly.

Automation bias and the deskilling signal

Then the second-order effect, which is the finding that should be changing how deployments are designed.

At four Polish endoscopy centres, after AI polyp detection entered routine practice, the adenoma detection rate in standard unassisted colonoscopy fell from 28.4 percent to 22.4 percent. The endoscopists were experienced, each with more than 2,000 procedures behind them. The study is observational and retrospective and should be read with those limits attached rather than as a verdict. But it is the first real-world evidence that exposure to assistance can degrade unassisted performance on an outcome patients feel, and it means the risk of a clinical AI tool does not end when the tool is switched off.

The liability sits with whoever signed

Both of the above land on a person whose name is on the record. The tool is a device or, increasingly, not even that. The signature is a licence. Until liability is allocated explicitly, the clinician is the default absorber of every failure mode above, which is why practical AI governance in healthcare is a workforce issue before it is a compliance one.

Risks the institution carries

Recall exposure and the evidence that was never published

Institutional risk starts at procurement, and it can be quantified.

A retrospective cohort of 903 FDA-authorised AI-enabled medical devices found 43 recalled, 4.8 percent, at a median 458 days from authorisation. Devices whose supporting clinical study information was missing had a higher hazard of recall than those with published studies. Devices flagged in postmarket surveillance carried a hazard ratio of 4.28. And incorrect use of the device, rather than a defect in the model, was implicated in 12 of the 31 devices with use-related problems, with an estimated hazard ratio of 3.33.

Read that last one twice. The most common trigger was not the algorithm failing. It was the algorithm being used in a way it was not validated for, which is an institutional control problem and shows up on the institution's balance sheet. Anyone building the business case should be reading it alongside what AI in healthcare costs, because a recall midway through a deployment writes off the integration work, not just the licence.

A new attack surface

The category missing from nearly every consumer-facing risks page is security, and it is the one where the institution has no shared defence to fall back on.

A 2026 Nature review of large language model safety and security in healthcare maps hazards to each stage of a clinical AI system's lifecycle, across design, data, model, inference and environment, and classifies threats by their current clinical relevance rather than their theoretical interest. The threat classes it covers include data poisoning during training, prompt injection reaching the model through content it processes, extraction of training data, and adversarial manipulation of inputs. These are not failures a clinician can catch by reading carefully, because a poisoned or injected system fails in exactly the confident register a working one uses. They are defended at the level of procurement, architecture and monitoring, or not at all.

Nobody is counting the incidents

The institutional blind spot is that there is almost no shared record of what has already gone wrong. Researchers who went looking found 15 candidate public repositories, of which 8 held health-related AI records, yielding 488 entries that deduplicated to just 295 unique incidents spanning 2012 to 2025, mostly from the US and UK. Their own reading of that number is the important part: 295 incidents across a decade of accelerating deployment reads as underreporting rather than as safety, because no mandatory channel exists to report into.

Aviation and pharmacovigilance both work because near misses are collected. Health AI currently has neither the channel nor the obligation, so every institution is learning from its own incidents alone.

The risks nobody has been assigned

What is left over is a short list, and it is the list that decides how the rest resolve.

Models drift after deployment as patients, coding practice and case mix change, so a system validated once is not validated permanently. Incidents are not systematically collected. Liability is undefined. An analysis that adapted the Swiss cheese model of safety to AI-driven digital health proposed five protective layers rather than a single control point: data governance and quality assurance, model development and change control, sociotechnical integration and human oversight, regulatory and ethical compliance, and post-market monitoring with incident response. It also makes the observation that separates AI failures from human ones. Human error is distributed and idiosyncratic. AI error is systematic, repeating consistently across similar cases, so at scale it can affect many patients in a very short window.

Regulators have started to move on precisely this. The UK convened a National Commission into the Regulation of AI in Healthcare and published findings drawn from public polling, deliberative research, an open call for evidence and the MHRA's AI Airlock programme in June 2026. That is the right direction and it is slower than deployment, which is the condition this whole field operates in.

What measurably reduces the risk

The listicles stop at the problem. The mitigation literature has numbers, and they are worth knowing before signing anything.

A systematic review of hallucination mitigation strategies found that retrieval-augmented generation grounded in high-quality medical sources reduced factual errors by 30 to 50 percent against base models. Knowledge graph integration cut hallucination rates by 25 to 45 percent, with the larger gains on fact-based rather than reasoning-based questions. Self-reflection and chain-of-verification frameworks improved factual accuracy by 20 to 35 percent. Human-in-the-loop review reached the highest overall accuracy at 85 to 95 percent, at a real cost in scalability.

Three findings held across the studies. No single mitigation layer was adequate alone. Data governance architecture, including provenance and source traceability, materially changed the risk. And effectiveness depended heavily on implementation context, particularly workflow burden and the escalation path for uncertain outputs. Combined approaches beat any single one, which is the same conclusion the safety-layer work reaches from the other direction, and it is the practical content of responsible AI in healthcare once the principles are set aside.

Why this does not get solved inside one department

Look at the four piles again. Undertriage is a clinical safety question. Contestability is a legal one. Recall exposure is procurement. Incident surveillance is regulation. Not one of them is answered by a better model, and not one of them is answered by a single profession working alone.

That is the reasoning behind running the Health Tech track alongside Government, AI Ethics and Finance at the EX Future Summit rather than in a separate room. A liability rule and a clinical safety threshold are the same conversation held by people who rarely share a table. The summit runs 18 to 20 November 2026 as a single continuous thirty-hour broadcast between Las Palmas and Bali, twelve hours apart, and online attendance is free for verified researchers and students.

FAQ

What is the biggest risk of AI in healthcare right now?

Not any single wrong output, but a wrong output that reaches a decision with no assigned owner. The sharpest measured case is consumer triage, where 51.6 percent of true emergencies in a controlled stress test were routed to a 24 to 48 hour follow-up rather than an emergency department. No clinician saw those recommendations, no institution logged them, and no regulator required the result to be reported.

Who is liable when AI causes harm to a patient?

In practice, the licensed human whose name is on the record. Formally, it is unsettled. Governance frameworks in use across healthcare organisations rarely address medicolegal liability, and only 13 percent give anyone a way to contest an AI-influenced decision at all. Liability allocation, not model accuracy, is the thing clinicians most often say would change their willingness to rely on these systems.

How often are AI medical devices recalled?

Across 903 FDA-authorised AI-enabled devices, 43 were recalled, or 4.8 percent, at a median 458 days after authorisation. Devices lacking published clinical study information carried a higher recall hazard, and the single most common recall trigger was use-related: the device being used outside the conditions it was validated for.

Is it safe to ask a chatbot about symptoms?

The measured failure pattern is specific and worth knowing. These tools are most reliable in the middle of the acuity range and least reliable at both ends, missing emergencies and escalating trivial complaints. Their recommendations also move when someone downplays the symptoms during the conversation. For anything that could be time-critical, the tool is the wrong first stop.

Do the disadvantages of AI in healthcare outweigh the benefits?

That framing does not survive contact with the evidence, because the risks are not evenly distributed across applications. Documentation and prediction at scale have the strongest support and the most contained failure modes. Autonomous clinical judgment, consumer-facing triage and any deployment without post-market monitoring carry risks that are currently unmanaged rather than merely present. The useful question is not whether to adopt but which layer of safeguard is missing from the specific deployment in front of you.

EX-AI-Summit 2026 · 18–20 November · Las Palmas (WET) · Bali (WITA) · Online
Presented by EX Venture Inc. · Seraph SL · Equation Labs SL

ProgramPartnersAboutContactLegalPrivacy

We use essential cookies to run this site and optional analytics cookies to understand how it is used. You can accept all or reject optional cookies. See our privacy notice and legal.