A sweeping review of artificial intelligence in American medicine has delivered a stark verdict: most AI and machine-learning devices cleared for patient care have never been tested to determine whether they actually improve patients’ health. Among 1,357 AI-enabled medical devices authorized by the U.S. Food and Drug Administration through December 5, 2025, researchers identified only three that had been evaluated using patient-centered outcomes such as survival, stroke, hospitalization, or quality of life. The analysis, published in PLOS Digital Health, suggests that the rapidly expanding medical-AI industry is advancing far faster than the clinical evidence needed to show that its products benefit the people expected to use them.
The devices examined in the study include technologies designed to support surgical planning, estimate cardiovascular risk, analyze mammograms and other medical images, and assist clinicians with a widening range of diagnostic and decision-making tasks. Many rely on machine-learning algorithms that identify statistical patterns in large datasets and use those patterns to classify images, generate risk scores, or recommend clinical actions. Yet technical performance is not the same as improved health. An algorithm can detect an abnormality accurately, for example, without proving that its use leads to earlier treatment, fewer complications, longer survival, or a better quality of life.
The researchers, led by Rawan Abulibdeh of the University of Toronto, systematically investigated how the FDA-cleared devices had been evaluated in human patients before entering clinical care. Their analysis found that only 34 of the 1,357 devices had been included in registered clinical trials. Results were publicly available for just 12 devices, while peer-reviewed manuscripts had been published for 12. The numbers indicate that even when developers conduct clinical research, the evidence may remain difficult for clinicians, regulators, and patients to access. More importantly, only three devices had undergone studies that measured outcomes directly meaningful to patients rather than focusing primarily on algorithmic accuracy or technical agreement with an existing tool.
The distinction is central to understanding the regulatory pathway. Many medical devices reach the U.S. market through a process in which manufacturers demonstrate “substantial equivalence” to an already authorized device. This pathway is intended to establish that a new product is sufficiently similar in safety and intended use to an existing one. For AI systems, however, similarity or technical performance does not necessarily establish clinical effectiveness. A newer algorithm may produce a result more quickly or match expert interpretations on selected cases, but those achievements do not automatically show that it reduces diagnostic errors in routine practice or improves outcomes across the diverse populations encountered in real healthcare systems.
The review also exposed serious limitations in the populations and settings represented by existing studies. Research was concentrated in highly resourced healthcare environments, while important patient groups were frequently excluded. Pregnant women, adults older than 75, and people who do not speak English were among the populations often missing from evaluations. These omissions matter because medical algorithms can behave differently when applied to patients whose age, language, disease severity, imaging equipment, treatment access, or demographic characteristics differ from those in the development dataset. A system that performs well in a specialized academic hospital may be less reliable in a rural clinic, an under-resourced health system, or a setting where patient records are incomplete.
The consequences extend beyond statistical uncertainty. If an AI system produces more false-negative results for a particular group, clinicians may miss serious disease. If it generates excessive false positives, patients may undergo unnecessary tests, procedures, anxiety, and expense. Such effects can be amplified when an algorithm becomes embedded in electronic health records or clinical workflows, where its recommendations may appear authoritative even when the underlying evidence is limited. The researchers warn that insufficient evaluation could allow AI tools to reinforce existing healthcare disparities rather than reduce them, particularly when institutions adopt systems without independently verifying their performance among local patients.
The authors also point to structural incentives that make rigorous outcome research difficult. Developers may face strong commercial pressure to bring products to market quickly, while long-term clinical trials can be expensive, time-consuming, and logistically complex. Demonstrating that a device changes hospitalization rates or quality of life requires larger and longer studies than showing that it detects a feature on an image. It may also require coordination across hospitals, careful monitoring for unintended effects, and methods capable of identifying differences between demographic and clinical subgroups. In this environment, technical validation can become the dominant benchmark, even though the ultimate purpose of medical technology is to help patients live longer, healthier, or more independent lives.
The evidence gap could also create international risks. Many countries, especially those with limited regulatory resources, look to decisions made by agencies in wealthier nations when determining whether to adopt new medical technologies. If a device has been cleared in the United States but has not been tested across diverse populations or healthcare environments, patients elsewhere may become inadvertent participants in a large, uncontrolled experiment. The concern is particularly acute in low- and middle-income countries, where differences in disease patterns, infrastructure, staffing, language, and access to follow-up care may alter how an AI system performs. A regulatory decision made in one setting can therefore influence clinical practice far beyond the population in which the device was developed.
To address these problems, Abulibdeh and colleagues propose redesigning the framework used to assess AI medical devices. Their suggested three-phase approach would move beyond a single premarket demonstration of similarity or technical performance and require evidence that systems are effective across diverse patient subgroups and healthcare settings. Such a framework would connect algorithm development with real-world clinical validation and continued evaluation after deployment. The approach reflects a broader shift in medical-AI research: from asking whether a machine can reproduce a label or prediction to asking whether its use changes decisions, improves care, avoids harm, and delivers benefits fairly.
The study does not claim that every AI medical device is ineffective, nor does it suggest that algorithms have no value in healthcare. Instead, it highlights how little is known about their effects on the outcomes patients care about most. As AI tools become increasingly visible in hospitals and clinics, clearance may be mistaken for proof of benefit. The researchers argue that these are fundamentally different claims. A device can be legally cleared because it resembles an existing product while still lacking evidence that it helps people live longer or better. Their findings place that distinction at the center of the debate over medical AI—and raise an urgent question for regulators, developers, clinicians, and patients: before intelligent systems become routine parts of care, what evidence should be required to show that they truly make care better?
Subject of Research: People
Article Title: 1,357 AI medical devices cleared, 3 actually tested on patient outcomes
News Publication Date: 19-Aug-2026
Web References: https://doi.org/10.1371/journal.pdig.0001597
References: Abulibdeh R, Cajas Ordóñez SA, Celi LA, Gorijavolu R, Izath N, Markussen Lunde T (2026). “1,357 AI medical devices cleared, 3 actually tested on patient outcomes.” PLOS Digital Health, 5(8): e0001597. DOI: 10.1371/journal.pdig.0001597
Image Credits: Abulibdeh R, Cajas Ordóñez SA, Celi LA, Gorijavolu R, Izath N, Markussen Lunde T, 2026, PLOS Digital Health, CC BY 4.0
Keywords: artificial intelligence, medical devices, machine learning, FDA clearance, clinical trials, patient outcomes, healthcare technology, medical imaging, health equity, digital health
Tags: AI in surgical planning and risk assessmentAI medical devicesclinical effectiveness of AI in medicineevaluation of AI diagnostic toolsevidence-based AI in clinical practiceFDA clearance of AI medical devicesimpact of AI on patient healthmachine learning algorithms in medical imagingpatient outcome testing in AI healthcarepatient-centered outcomes in medical AIregulatory gaps in AI healthcare devicessafety and efficacy of AI medical technologies





