Summit
AI in Healthcare: What the 2026 Evidence Shows
Adoption numbers, what holds up outside the study, the deskilling signal, and who actually pays: an evidence map of AI in healthcare in 2026.

Most writing about AI in healthcare is still a catalogue of what the technology could do. That catalogue was useful when the answer was speculative. It is not useful now, because the systems are already installed, and the interesting question has moved: of everything being claimed, which parts survive contact with a Tuesday morning clinic?
The 2026 evidence base is finally large enough to answer that, and it does not sort neatly into good news and bad news. It sorts into four piles. Some applications are proven. Some are impressive in evaluation and fragile in practice. Some are unevenly distributed to the point of becoming an equity problem. And some are held up by nothing technical at all.
How much is actually in use
Start with deployment, because most market commentary substitutes a growth forecast for a measurement.
In US hospitals, 71 percent reported using predictive AI integrated with their electronic health record in 2024, up from 66 percent in 2023. That is risk scoring: readmission likelihood, early deterioration, appointment no-shows. It is the oldest layer and by some distance the most embedded.
Physicians report something similar from the other side of the desk. In the American Medical Association's 2026 survey of nearly 1,700 doctors, 81 percent reported awareness or use of AI, double the share in 2023, with the average number of use cases rising from 1.1 to 2.3. The applications named most often are unglamorous: summaries of medical research and standards of care at 39 percent, discharge instructions and care plans at 30 percent, documentation at 28 percent. Assistive diagnosis, the thing that gets written about, sits at 17 percent.
An independent survey run by Doximity across 3,151 physicians in 15 specialties puts current use at 54 percent and, more usefully, captures the slope: adoption rose from 47 percent in its spring 2025 wave to 63 percent in the wave closing January 2026.
Generative AI specifically is much further back than the headlines suggest. Analysis of the same 2024 hospital survey found 31.5 percent of hospitals were early adopters of generative AI integrated with the EHR, against 43.7 percent classed as delayed adopters. Predictive AI is infrastructure. Generative AI is still a pilot in most buildings.
Where the evidence holds up
The January 2026 State of Clinical AI report, produced by researchers across Stanford, Harvard and affiliated health systems, reviewed the most influential clinical AI studies of 2025 with one question in mind: what survives outside a controlled setting. Its answer has a clean organising principle, which is that these systems perform best when they address problems where humans are limited by scale rather than judgment.
Prediction at scale
This is the strongest tier. A model trained on continuous wearable vital signs predicted patient deterioration 8 to 24 hours before standard hospital alerts, identifying people at risk of ICU transfer, cardiac arrest or death while intervention was still possible. Another estimated biological age from routine health records across millions of individuals and predicted mortality more accurately than epigenetic clocks or frailty scores. Neither task requires clinical judgment. Both require reading more data than a person can hold.
Documentation, and the hours it gives back
Ambient note generation reached broad adoption in 2025, and the effect sizes are unusually large for health IT: across multiple hospital systems, physicians reported spending up to 83 percent less time writing notes, with one system reporting a 112 percent return on investment. This is the least discussed and most consequential category, and it is why so much of the money flows to AI in healthcare administration rather than to the clinical frontier.
Assisted reading, not autonomous reading
The pattern repeats wherever a clinician stays in the loop. In Germany, radiologists who could optionally consult an AI system detected more breast cancers without increasing false alarms. In Kenya, a background review system deployed with Penda Health reduced diagnostic and treatment errors across tens of thousands of urgent care visits. The consistent finding across settings is that AI supporting a clinician beats AI replacing one, and that how a tool is introduced matters as much as how it performs.
Where it breaks down outside the study
Now the second pile, which the vendor explainers skip entirely. The same State of Clinical AI review supplies most of it.
Reasoning collapses when the question stops being a quiz
Researchers took standard medical multiple-choice questions and changed the correct answer to "none of the other answers." The clinical reasoning required did not change. Accuracy dropped across leading systems, in some cases by more than a third. Performance fell again when models had to ask follow-up questions, work with incomplete information, or revise a decision as details emerged. On tests of reasoning under uncertainty, systems performed closer to medical students than to experienced physicians, and tended to commit hard to an answer precisely when ambiguity was highest.
Hold that against the most quoted result of the year, recorded in Stanford's AI Index: a multi-agent system paired with a frontier model scored 85.5 percent on complex published case studies, against 20 percent for physicians working without their usual tools. Both findings are real. Only one of them describes a clinic.
Most studies do not resemble medicine
The structural version of the problem: a review of more than 500 medical AI studies found nearly half tested models on exam-style questions, and only 5 percent used real patient data. Very few measured whether a model recognised its own uncertainty. Fewer examined bias.
Clinicians do not spend their day answering board questions. They spend it reviewing charts, triaging messages, coordinating care and deciding when not to intervene. An evaluation regime built on quizzes cannot tell you how a tool behaves in that work, which is a large part of why the risks of AI in healthcare are still argued about rather than measured.
The deskilling signal
The finding that should change how deployments are designed came from four endoscopy centres in Poland. After AI polyp detection was introduced into routine practice, the adenoma detection rate in standard, non-AI assisted colonoscopy fell from 28.4 percent to 22.4 percent, an absolute drop of 6 percentage points, in a multicentre study published in The Lancet Gastroenterology and Hepatology. The endoscopists were experienced, each with more than 2,000 procedures behind them.
The study is observational and retrospective, and it should be read with those limits attached rather than as a verdict. But it is the first real-world evidence that AI exposure can degrade unassisted clinical performance on an outcome that matters to patients, and it complicates every trial that used unassisted clinicians as its control group.
The adoption gap inside a single country
National averages hide the more urgent story. Within the same 71 percent figure, 86 percent of hospitals affiliated with multi-hospital systems reported using predictive AI against 37 percent of independent facilities. Urban hospitals ran at 81 percent, rural at 56 percent. Critical access hospitals, the small facilities at least 35 miles from the next one, sat at 50 percent against 80 percent elsewhere.
Governance follows the same money. Among hospitals using these models, 82 percent evaluated them for accuracy and 74 percent for bias. Evaluation capacity is expensive, and the institutions least able to afford it are also the ones least able to absorb a bad model. This is a digital divide forming in real time, not a lagging tail that catches up on its own.
The rulebook is 240 documents, not one
Regulatory clearance is routinely mistaken for evidence of benefit. The distinction is worth stating precisely. By the AI Index count, the FDA authorised 258 AI medical devices in 2025, and the State of Clinical AI report puts the cumulative total of cleared AI-enabled tools above 1,200. Most entered through device-modification pathways that rely on existing safety and efficacy evidence rather than new trials. Among devices with clinical studies, only 2.4 percent were supported by randomised trial data.
Around that sits a policy layer that is large and thin at the same time. A January 2026 snapshot of the Health and AI Policy Index holds 240 policies relevant to AI in health care, of which just 22, or 9 percent, are rated high impact, with transparency and governance tags on 144 of them. There is no single omnibus health-AI law in the United States and no single regulator. What exists is volume without settlement, which is a familiar problem to anyone tracking sovereign AI and who controls the stack.
Physicians have priced this accurately. In the AMA survey, clear liability frameworks ranked highest among the regulatory actions that would build trust, ahead of any technical assurance. That is a demand for AI governance in healthcare that resolves who is answerable, not for more guidance documents.
Nobody has settled who pays
The least covered constraint is also the one currently deciding which tools get built.
As of January 2026 there were 26 CPT codes for clinical AI solutions, only three of which hold Category I status. Most of the rest are Category III and priced case by case by regional Medicare contractors, so payment varies by geography and cannot be planned around. No public database tracks AI-specific billing at all. The predictable consequence is that capital flows to administrative AI, where the return does not depend on a coverage decision.
The Peterson Health Technology Institute put the structural version of this to health system leaders and found that fee-for-service, pay-for-performance and capitation all misprice increasingly autonomous clinical AI. Fee-for-service pays for clinician time when clinician time is no longer the input creating the value. The CMS ACCESS model, which replaces per-service billing with outcome-aligned payments, is described in that work as the most concrete step so far, though analysts have called its rates of 180 to 420 dollars per beneficiary per year financially unviable. Anyone modelling what AI in healthcare actually costs is modelling an unfinished payment system.
What actually decides the next two years
Not model capability. On the evidence above, the three unresolved variables are evaluation that reflects real workflows, a liability position clinicians can rely on, and a payment route that does not punish clinical applications for being clinical. Each is a policy and institutional problem rather than a research one, which is inconvenient for everybody currently working on the research.
That is also why these arguments do not stay in one department, and why the EX Future Summit runs its Health Tech track alongside Government, Finance and AI Ethics rather than in a separate room. A reimbursement rule and a clinical safety threshold are the same conversation held by people who rarely share a table, a point that recurs in the debate over the ethics of AI in healthcare. The summit runs from 18 to 20 November 2026 as a single continuous thirty hour broadcast between Las Palmas and Bali, twelve hours apart, and online attendance is free for verified researchers and students.
FAQ
Will AI replace doctors?
Nothing in the current evidence points that way, and the strongest results point the other way. Randomised trials found physicians using AI alongside standard resources made better treatment decisions than those using conventional tools alone, while unassisted performance is the thing that shows signs of degrading. The reliable configuration is clinician plus system, and it has been consistent across imaging, primary care and urgent care settings.
What is the difference between predictive AI and generative AI in healthcare?
Predictive AI scores risk from structured data: who is likely to deteriorate, be readmitted, or miss an appointment. Generative AI produces text, such as a note, a summary or a draft reply. The gap in maturity is large. Roughly seven in ten US hospitals run EHR-integrated predictive AI, while under a third are early adopters of generative AI, and hospitals already using predictive models are far likelier to adopt generative ones.
Are AI health tools approved by the FDA?
Some are, as regulated medical devices, and many are not devices at all. Clearance also carries less evidentiary weight than the word suggests: most authorisations proceed through modification pathways that reuse existing evidence rather than requiring new trials, so a cleared tool has met a safety and equivalence bar, not a proof of clinical benefit in your population.
Should patients trust AI answers about their own health?
The realistic framing is that avoidance is no longer an option: AI-generated summaries now appear at the top of 84 to 92 percent of health-related searches, rising to 92 percent for symptom queries. Physicians surveyed were broadly comfortable with patients using AI for general health and medication questions, and markedly less so where clinical judgment applies. Nearly half strongly opposed patients using AI to interpret their own radiology or pathology results.
What is the biggest barrier to hospitals adopting AI?
Not accuracy, though 71 percent of surveyed physicians name accuracy and reliability as their top concern. The binding constraints are structural: no dependable reimbursement route, unresolved liability, and evaluation capacity that independent, rural and critical access hospitals largely do not have. Those three explain the adoption gap better than any measure of model quality.