Uncategorized

Only 3 of 1,357 FDA-Cleared AI Medical Devices Have Been Tested on Whether They Actually Help Patients

A University of Toronto-led study finds that of 1,357 FDA-cleared AI medical devices, only three have ever been tested on whether they improve patient outcomes like survival or hospitalization.

aidatanews

A new analysis published August 19 in PLOS Digital Health has delivered one of the starkest reality checks yet on the AI-in-medicine boom: of the 1,357 artificial intelligence and machine-learning medical devices the U.S. Food and Drug Administration has cleared for use, only three have been evaluated for whether they actually improve patient outcomes like survival, hospitalization, or quality of life. The study, led by Rawan Abulibdeh of the University of Toronto alongside collaborators including Leo Anthony Celi of MIT Critical Data, systematically cross-referenced every FDA-cleared AI device against ClinicalTrials.gov and PubMed through December 5, 2025.

The numbers behind the gap

The researchers found that just 2.5% of the 1,357 cleared devices were linked to a registered, prospective clinical trial. Fewer still — 0.9%, or 12 studies — had posted results, and the same share had produced a peer-reviewed publication. Only 0.2%, three studies in total, actually measured patient-centered outcomes such as mortality, morbidity, or hospital readmission. The vast majority of clearances instead relied on retrospective, technical-performance benchmarks — comparing an algorithm’s output to a reference standard — rather than tracking what happens to real patients treated with AI assistance.

Why the FDA pathway allows this

Most AI medical devices reach the market through the FDA’s 510(k) pathway, a route designed for products that are “substantially equivalent” to an already-cleared device or predicate. That pathway does not require the kind of large, randomized outcome trials that new drugs typically undergo. The FDA’s AI-Enabled Medical Device List has ballooned from a handful of entries in the mid-2010s to more than 1,500 by early 2026, a growth curve driven overwhelmingly by radiology and imaging tools that can point to strong technical accuracy metrics without ever being tracked into a hospital’s actual patient outcomes.

Who gets left out of the evidence

The study also flagged a representativeness problem layered on top of the evidence gap: most of the handful of studies that did exist were conducted in highly resourced healthcare systems, and they frequently excluded groups such as pregnant patients, adults over 75, and non-English speakers. That means even the thin slice of AI tools with published clinical data may not generalize to the populations most likely to be affected by errors — older patients with more comorbidities, or patients treated in under-resourced hospitals with different equipment and staffing than the academic medical centers where many AI tools are validated.

The researchers’ own framing

Abulibdeh and colleagues wrote that “the market is crowded with AI tools, yet the evidence supporting patient-centered outcomes remains remarkably thin,” describing what they called an “evidence attrition” problem — each step from initial development to regulatory clearance to real-world outcome validation loses the vast majority of candidate devices, leaving only a handful that have ever proven they change what happens to a patient rather than simply matching a radiologist’s read on a test set.

Industry pushback and a more measured view

AI device makers and some clinicians argue the comparison to drug trials is unfair: unlike a pill, an imaging-triage algorithm is typically a decision-support tool layered on top of a human clinician’s judgment, not a standalone therapy, so traditional outcome trials may be slower and costlier than the risk profile justifies. Proponents also note that dozens of AI tools have been in routine clinical use for years without documented harm, and that demanding outcome trials for every 510(k) clearance could slow beneficial tools from reaching hospitals. Critics counter that the absence of harm data is not the same as evidence of benefit, and that regulators and hospitals have effectively been running an uncontrolled real-world experiment on patients without the monitoring infrastructure to know if it is working.

What happens next

The study’s authors are calling for the FDA to require post-market surveillance data and outcome tracking as a condition of continued clearance, rather than treating clearance as a one-time technical checkpoint. With the FDA’s breakthrough-device pipeline reportedly filling with generative and foundation-model-based tools through 2026, the pressure to close this evidence gap is likely to intensify rather than fade — and hospitals purchasing AI software will increasingly be asked to answer a question regulators largely have not: does this tool make patients better off, not just faster diagnosed.