Summit
The Future of AI in Healthcare: What Is Already Scheduled
Not forecasts. The future of AI in healthcare read off what is already dated: an FDA GenAI docket, 15 to 20 Phase III readouts, and a validation gap.

Almost everything written about the future of AI in healthcare is a forecast. The competent versions are worth reading: BCG's 2026 outlook describes agents that observe, plan and act on their own, co-pilots that absorb documentation, and health systems deploying models to predict and prevent illness. That is a serious argument by serious people. It is also, as a genre, unfalsifiable inside the window a reader cares about.
There is a second document available, and hardly anyone writes it. Parts of the next two years are already fixed. A federal comment docket closes on a specific date. A defined number of pivotal drug trials report. A published count exists of how many deployed AI tools have ever been tested on a patient outcome. None of that is speculation, and all of it will grade the forecasts.
This is that second document.
The gap every forecast has to cross first
Start with the number that constrains everything else.
A systematic analysis published in PLOS Digital Health examined every AI and machine learning enabled medical device cleared or approved by the US FDA through 5 December 2025. Of 1,357 cleared devices, 34 (2.5%) were linked to a registered prospective trial, 12 (0.9%) reached peer-reviewed publication, and 3 (0.2%) were evaluated for patient-centered outcomes such as mortality, morbidity or readmissions. Three. A separate review of 691 cleared devices found 1.6% cited randomised trial data at all.
The reason is structural rather than scandalous. Nearly all of these devices reached the market through the 510(k) pathway, where a sponsor demonstrates substantial equivalence to an already legally marketed predicate device. That is a claim about resemblance, not about benefit. Evidence gaps propagate down a chain of predicates, and the chain can start outside AI entirely: a review of 510(k) clearances found that up to one-third of AI and machine learning devices were cleared on the basis of predicates that were not themselves AI devices.
So the honest reading of the installed base is that regulatory authorisation has run well ahead of clinical validation. Any forecast about what AI does for patients in 2028 is a forecast about closing that gap, whether or not it says so. For the fuller picture of what the 2026 evidence base actually shows, the deployment numbers are a separate matter from the outcome numbers.
The part that already works, and the reason it works
One application has escaped this problem, and it is worth understanding exactly why, because it is the template.
Ambient documentation now has multi-site quantitative evidence. A longitudinal cohort study across five US academic health systems compared 1,809 AI scribe adopters against a total sample of 8,581 clinicians and found 13.4 fewer minutes of electronic health record time, 16.0 fewer minutes of documentation time, and 0.49 additional weekly visits delivered. A quality improvement study across six health systems found burnout among 263 ambulatory clinicians fell from 51.9% to 38.8% after thirty days. Outside the US, a sixteen-month deployment in Spanish outpatient care logged 2,339,281 assisted encounters, adoption rising from 2.72% to above 30% of visits, and semantic agreement holding between 87.4% and 89.2%.
Two things about that record. The effects are modest, not transformative: thirteen minutes, half a visit a week, and in the JAMA cohort no significant change in after-hours EHR time at all. And the mechanism is unglamorous. The model drafts, the clinician edits, and every word is read by a licensed human before it enters the record. The failure mode is caught by somebody who was already standing there.
That is the design pattern. The applications that promise more are the ones removing that reader, and removing the reader is precisely what the next section is about.
Generative AI reaches the clinic through a docket, not a launch
The most consequential scheduled event in this field is a comment deadline, which is why it appears in no listicle.
On 18 August 2026 the FDA's Center for Devices and Radiological Health, through its Digital Health Center of Excellence, issued a discussion paper on the regulation of generative AI enabled medical devices and opened docket FDA-2026-N-7874, with comments due by 19 October 2026. It is not draft guidance and it changes nothing today. Several trade outlets reported otherwise on release day. What it does is describe the architecture the agency expects to build, and the record it collects in those sixty-two days is what the architecture will be built from.
Two axes and a credentialing exam
The proposed risk model has two dimensions: how independently a device function acts, running from non-directive information through action-directing recommendations to supervised and then fully autonomous action, and the severity of harm if the output is wrong. Where a product lands on that grid sets its evidence burden.
The premarket approach is the genuinely novel part. Rather than testing every possible input, the paper proposes competency assessment modelled on how physicians are credentialed: non-clinical device benchmarking against standardised datasets, followed by clinical confirmation, with explicit relief that clinical confirmation might not require a prospective clinical study in every case. Board exams, then residency, then maintenance of certification. The same document floats a voluntary Foundation Model Device Master File, so a model developer could file architecture and guardrail detail confidentially and a device sponsor could reference it, and it poses twenty-six open questions rather than answering them.
How far the frontier has actually moved
For calibration on where the market sits today: only one cleared device, UpDoc V1.0, incorporates a patient-facing large language model, and it cleared against a drug-dose calculator predicate with the language model kept outside the insulin dosing logic. The model does intake and communication. Deterministic code makes the decision that matters. No cleared device claims autonomous diagnosis.
The distance between that and an agentic system with clinical authority is the distance the discussion paper is trying to survey. The usual cycle from discussion paper to proposed rule runs roughly eighteen to thirty-six months, which puts a first binding framework somewhere in 2028. Teams building on generative AI in healthcare are operating in that interval, and so is anyone doing AI regulatory compliance work on a product that ships before the rules exist.
The drug pipeline gets graded between now and 2027
The other scheduled reckoning is in pharmaceuticals, where the claim has been outstanding for a decade and the exam finally sits.
Roughly 175 AI-discovered programmes are in clinical development with zero regulatory approvals, and between 15 and 20 reach pivotal Phase III readouts through 2026 and 2027. The BCG analysis of AI-discovered molecules is the one to hold onto, because it splits cleanly. In Phase I, success runs at 80 to 90% against a historical norm near 50%. In Phase II, success falls back to roughly 40%, the ordinary industry rate.
The split is explicable and it is the whole story. Phase I failure is chemistry: toxicity and pharmacokinetics are molecular properties with decades of training data behind them, and generative design optimises for exactly that. Phase II failure is biology: the target turns out not to matter for the disease, or the effect is too small in real patients. Nothing in the molecule fixes a wrong hypothesis about the disease.
Two caveats before anyone uses those figures. Other analyses put AI Phase II success in the 65 to 75% range from smaller and more selectively reported samples, so the sources genuinely disagree and the conservative read is the defensible one. And the readouts are not equivalent. Rentosertib dosed its first patients in a 320-person Phase III on 7 July 2026 against a target the AI itself selected. Zasocitinib has already cleared Phase III, but against TYK2, a target validated by an approved drug years earlier. A win on the first is a categorically larger result than a win on the second, and press coverage will not make the distinction.
Judge the set by its base rate. Roughly half of all Phase III candidates fail regardless of how they were discovered. One success out of twenty readouts proves nothing. Six would be a real signal.
Three constraints that will set the pace, none of them technical
Nobody has agreed who pays
Medicare has assigned payment to roughly ten of the 1,524 devices on the FDA list. A hospital deploying a cleared tool is usually absorbing the cost against an efficiency argument, and efficiency arguments do not survive a bad budget year. This is the least discussed and most binding constraint on adoption, and it is treated at length in our note on who pays for clinical AI.
The map of what gets built is lopsided
A cross-sectional analysis of all 1,430 FDA authorisations between September 1995 and December 2025 found radiology accounted for 1,094 of them (76.5%), the top three panels for 90.6%, and several enormous clinical specialties for almost nothing: pathology 9 devices, microbiology 6, obstetrics and gynaecology 4, and no authorisations at all under a psychiatry or behavioural health review panel. Annual volume rose from a mean of 1.8 per year between 1995 and 2014 to 331 in 2025, and 502 of 740 companies hold exactly one device.
Read that as a map of where the labelled images were, not of where the clinical need is. Any forecast that describes AI transforming mental health care is describing a category with no regulatory footprint whatsoever.
Validation is happening in the wrong places
The PLOS authors flag something that rarely reaches Western coverage: FDA clearance often functions as a gateway for global deployment, so under-validated tools enter low and middle income health systems without contextual validation. A model trained on one population and cleared in another is being asked to generalise twice.
The constructive answer in the literature is anticipatory regulation rather than static rules: bounded environments where a tool is tested with the regulator present. The MHRA AI Airlock, Singapore's IMDA sandbox and Article 57 of the EU AI Act are the working examples. Smaller jurisdictions have a structural advantage here, which is the same argument behind sovereign AI capacity and behind the Bali Province digital residency sandbox, which opens a cohort with a ninety-day residency, a policy pilot lane and access to municipal data. An island is a testable population with one regulator.
What to watch, with dates
Five items, all checkable against a public record.
- 19 October 2026. The FDA docket closes. Who filed, and whether the evidence burden shifts from premarket to postmarket, tells you the shape of the next framework.
- Through 2027. The Phase III readouts. Count the denominator, check whether the target was novel or already validated, and hold the results against a 50% base rate.
- The first LLM inside a decision. Today the one cleared example keeps the model outside the dosing logic. The clearance that changes that is the real threshold, not a product launch.
- Payment coverage. If Medicare assignment moves off single digits, adoption economics change. If it does not, deployment stays a cost centre.
- The specialty mix. Radiology has held above three quarters of authorisations in every independent count since 2021. Movement there means the data problem is being solved somewhere new.
None of this settles whether AI improves medicine. It sets out what would count as evidence, and when. That distinction is the substance of the governance of AI in healthcare argument, and it is the reason the summit runs Health Tech, AI Ethics and Global Financial Raise as separate tracks in the same building: the researcher generating the evidence, the regulator deciding what counts, and the capital deciding what gets built are usually three conversations in three cities. The summit runs a continuous thirty-hour broadcast between Las Palmas and Bali, free online for verified researchers and students, so at least the timezone is not the obstacle.
FAQ
Will AI replace doctors?
Not on the current regulatory record. No cleared device claims autonomous diagnosis, and the single cleared device with a patient-facing language model deliberately keeps that model outside the clinical decision. The FDA's own proposed risk grid treats fully autonomous action as the far end of a scale that nothing on the market has reached.
How many AI medical devices has the FDA approved?
The public list held 1,430 authorisations between September 1995 and December 2025, with 331 in 2025 alone. Approved is the wrong word for most of them. Nearly all arrived through the 510(k) clearance pathway, which asks for equivalence to an existing device rather than proof of benefit.
Has any AI-discovered drug been approved?
No. Roughly 175 programmes are in clinical development and none has reached approval. The first pivotal Phase III readouts arrive through 2026 and 2027, which is when the question becomes answerable rather than arguable.
What is the biggest obstacle to AI in healthcare over the next five years?
Evidence and payment, not capability. Three of 1,357 cleared devices have been evaluated for patient outcomes, and Medicare has assigned payment to roughly ten. Neither problem is technical, and neither is solved by a better model.
How can a country regulate medical AI before the evidence exists?
With anticipatory instruments rather than static rules. A regulatory sandbox lets a tool be tested in a bounded environment with the regulator in the room, generating the evidence and the rule at the same time. The MHRA AI Airlock, Singapore's IMDA sandbox and Article 57 of the EU AI Act are the working precedents.