Medical AI Has a Proof Problem

Published on: Sep 18, 2026
Author: Nigel Trimmer

If a machine can name the disease but still cannot be trusted to find it, what exactly have we bought? That is the awkward question hanging over medical AI: dazzling in the abstract, fragile in the clinic. The latest evidence does not suggest the technology is useless. It suggests something more uncomfortable for investors, hospitals, and vendors alike: competence in a controlled test is not the same as usefulness in the open world, where patient data arrive incomplete, messy, and misleading, like weather rather than geometry.

The temptation is old. We have seen it before in finance, where a model works beautifully until the regime changes, and in war, where a plan survives only until contact. Medicine is no different. A system can be trained to perform a neat trick in a narrow setting and still fail the real trial of practice. The market loves the clean curve, the polished demo, the language of inevitability. Nature does not care. It rewards systems that absorb shocks, not those that merely look elegant under glass.

The test that should unsettle everyone

A study published in JAMA Network Open on April 13 tested 21 AI models on 29 standardized clinical cases. According to reporting on the study, failure rates exceeded 80% for differential diagnosis when patient data were incomplete. That is the point where medicine begins, not ends: a doctor rarely receives a complete and tidy file from the world. The same models did far better when the data were complete. Failure rates fell below 40% for final diagnoses, and the top models exceeded 90% accuracy. That is impressive, but it also reveals the trap. A machine that performs best after the puzzle is nearly solved is not yet a reliable puzzle solver.

The lead author, Arya Rao of Mass General Brigham, put the distinction plainly: “These models are great at naming a final diagnosis once the data is complete, but they struggle at the open-ended start of a case, when there isn’t much information.” That is not a small caveat. It is the whole game. The first stage of diagnosis is not a finishing contest. It is a search problem under uncertainty, the kind of environment where human clinicians earn their keep and where brittle systems are exposed.

Why the benchmark can mislead

This is the familiar engineering error of mistaking a load test for real use. A bridge that survives a static weight in a yard is not therefore safe in a storm, on shifting ground, with fatigue, ice, and traffic all in play. In the same way, a medical AI that clears a narrow benchmark may still stumble when symptoms conflict, histories are partial, or the patient is not a textbook case. That is not failure in the abstract. It is failure at the edge, where harm lives.

The models in the JAMA Network Open study included systems from OpenAI, Google, Anthropic, xAI, and DeepSeek. That mix matters less as a leaderboard and more as a warning: this is not a single vendor’s problem. It is structural. When many models share the same weakness, the issue is not branding. It is the shape of the task. Closed-form questions favor pattern completion. Real medicine demands judgment under ambiguity, a less glamorous skill and one that resists easy scaling.

The evidence gap is wider than the hype

Even that may not be the largest problem. The deeper fragility lies in how little of the current AI medical device ecosystem has been tied to patient outcomes. A PLOS Digital Health study reported by AuntMinnie found that only 3 of 1,357 FDA-cleared AI medical devices, or 0.2%, had been tested on patient-centered outcomes such as mortality or readmissions. In radiology, only 3 of 1,059 cleared AI devices, or 0.3%, had registered prospective trials. Those are not rounding errors. They are signs of an industry that has built a large inventory of clearance without a comparable inventory of proof.

The same study found that 62% of AI device studies used small, homogenous cohorts and frequently excluded pregnant women, children, and non-English speakers. In other words, the evidence base does not just look thin. It looks narrow in the exact places where medicine should be broad. This is a classic institutional mistake: optimize for what is easiest to count, then call the result representative. But a sample that leaves out the difficult cases is like a map that omits the mountains and then claims to show the terrain.

Clearance is not the same as benefit

Sebastián A. Cajas Ordóñez of the MIT Critical Care team captured the problem from another angle: “We expected the evidence base to be thin. We did not expect three.” He was referring to the tiny number of devices with patient-centered outcome testing. He also said, “Under the 510(k) pathway, it clearance means a device is substantially equivalent to something already on the market, not that using it makes patients better off.” That is the distinction many buyers prefer not to dwell on. Clearance can mean sameness. It does not automatically mean improvement.

That distinction matters because medicine is not a place where “good enough” can be judged by software metrics alone. A model may reduce clerical burden, speed triage, or help organize information. Those are real benefits. But if the claim is that it improves care, then the burden of proof should move from engineering convenience to clinical outcome. Otherwise the field risks the oldest con in modern systems: confusing activity with value. A dashboard can glow while the underlying machine keeps overheating.

The psychology of believing too soon

Investors are especially vulnerable to this story because it has the shape of inevitability. AI is advanced, medicine is vast, and a large market is waiting. That is enough to seduce the forward-looking mind. But markets often overpay for narratives that compress complexity into a single trend line. The human brain loves extrapolation because it is cheaper than skepticism. Yet probability punishes cheap thinking. In a domain with rare but severe errors, the average outcome matters less than tail risk. One confident mistake can outweigh a hundred correct suggestions.

Medical AI is therefore not just a software story. It is a governance story. The question is not whether the model sounds intelligent. The question is whether the institution using it can measure harm, detect drift, and resist the tendency to accept performance claims as proof of benefit. In game theory, this is the difference between a signal and a credible signal. Anyone can announce excellence. Fewer can demonstrate it under conditions that cannot be gamed.

What an evidence ladder would change

The MIT Critical Data team proposes a three-phase evidence ladder: Phase 0 retrospective validation, Phase 1 prospective studies of at least 500 patients, and Phase 2 multi-center outcome trials of at least 2,000 patients. The team also calls for FDA pre-registration of prospective trials for Class II and III AI devices. That is not bureaucracy for its own sake. It is an attempt to prevent the field from declaring victory in the laboratory and calling it medicine.

A proper ladder of evidence is not anti-innovation. It is what makes innovation durable. In engineering, a design earns trust by surviving load, fatigue, and time. In biology, a trait survives because it fits an environment that keeps changing. In finance, the best systems are not those that predict every move but those that can remain solvent when they are wrong. Medical AI should be held to a similar standard. The goal is not to make the model look wise. The goal is to make the system safer when the model is uncertain.

The real test is still ahead

There is a temptation to say the technology will simply improve and the problem will vanish. Maybe it will. But markets and institutions should not invest as if a future fix is already here. A system that performs well only when fed complete data is still dependent on ideal conditions. And ideal conditions are a luxury medicine never has. Patients arrive late, confused, anxious, or in pain. Records are incomplete. Symptoms overlap. Language, age, and access shape what is seen and what is missed.

That is why the current evidence matters more than the promotional arc. It shows that medical AI can be strong in a narrow, tidy environment and weak at the opening move. It also shows that the validation pipeline is still thin where it counts most: in patient outcomes, in diverse cohorts, and in prospective trials that resemble real care. The lesson is not to reject the machine. It is to stop mistaking fluent output for clinical truth. In medicine, as in markets, the cost of early confidence is usually paid by those least able to absorb it.

AI Healthcare Services