Regulatory pathway for clearing an AI-based radiology diagnostic algorithm — predicate comparison, bias-tested validation, and FDA review
Before a single validation study is designed, a sponsor must decide which regulatory pathway applies. For the vast majority of AI-based radiology software, that decision hinges on finding a "predicate" — an already-legally-marketed device with the same intended use and similar technological characteristics — which permits clearance via the 510(k) premarket notification pathway rather than the far longer De Novo or Premarket Approval (PMA) routes.
The 510(k) pathway does not require a sponsor to prove independent safety and effectiveness from scratch, the way a PMA does. Instead, the sponsor must demonstrate that its new device is "substantially equivalent" to a legally marketed predicate: it has the same intended use, and either the same technological characteristics, or different characteristics that do not raise new questions of safety and effectiveness and that are supported by data showing the new device is at least as safe and effective as the predicate.
For an AI radiology triage or detection algorithm, this typically means finding a predicate that reads the same imaging modality (CT, MRI, X-ray, mammography), flags the same finding (e.g., pulmonary nodules, intracranial hemorrhage, bone fractures), and is used by the same class of clinician in the same clinical workflow. The predicate does not need to be recently cleared — devices cleared many years earlier remain valid predicates, and a chain of prior 510(k)s (sometimes called "predicate creep") has let each new AI generation cite the previous AI generation as its comparator.
Over 97% of FDA-authorized AI/ML-enabled medical devices have been cleared through the 510(k) pathway rather than De Novo or PMA — a reflection of how mature and well-precedented the AI-in-radiology device category has become since early computer-aided detection (CAD) clearances decades ago.
The De Novo pathway exists for novel, low-to-moderate-risk devices that have no valid predicate — it establishes a brand-new device classification and, once granted, itself becomes a predicate that later devices can cite. The PMA pathway is reserved for higher-risk (Class III) devices and requires standalone clinical evidence of safety and effectiveness, closer in rigor to a drug approval.
Both alternative pathways take substantially longer and cost substantially more than a 510(k). Because so many AI radiology use cases now have an established predicate lineage, sponsors are strongly incentivized to design their algorithm's intended use statement to fit within an existing device classification wherever legitimately possible — the predicate search is therefore often one of the earliest and most consequential regulatory-strategy decisions a company makes, well before the validation study is even designed.
An AI model is only as trustworthy as the data used to evaluate it. FDA expects a validation dataset that is genuinely independent of model training data, large and diverse enough to reflect real-world use, and analyzed not just in aggregate but broken out across demographic and equipment subgroups — because an algorithm with strong average performance can still fail specific populations.
A model evaluated on data it has already seen — even indirectly, through patients, scanners, or sites shared with the training set — will report inflated performance that will not hold up in the field. FDA reviewers scrutinize how a sponsor partitioned its data: ideally by patient and by clinical site, with the validation set drawn from institutions, scanner models, and patient populations distinct from training.
This "distribution shift" problem is especially acute for imaging AI, where subtle differences in scanner manufacturer, acquisition protocol, slice thickness, or contrast timing can measurably change model output. A validation set that only reflects the sites used to build the training data will systematically overstate real-world generalizability.
Aggregate sensitivity and specificity numbers can mask meaningful disparities. FDA increasingly expects sponsors to stratify performance by age, sex, race/ethnicity, body habitus, and — uniquely relevant to imaging — scanner manufacturer and model, since AI models can be surprisingly sensitive to acquisition hardware and reconstruction algorithms rather than the underlying pathology alone.
Published evaluations of medical imaging AI have documented subgroup performance gaps of 10–20 percentage points when training data lacked adequate representation of certain patient populations. A validation plan that cannot demonstrate consistent performance across the subgroups likely to use the device in practice invites an Additional Information request or, in severe cases, a refusal to clear.
FDA's 2021 Artificial Intelligence/Machine Learning (AI/ML)-Based Software as a Medical Device Action Plan explicitly identified data diversity and algorithmic bias as a top-priority risk area, pushing sponsors toward mandatory subgroup performance reporting rather than aggregate metrics alone.
With a curated, bias-tested dataset in hand, the sponsor runs a formal clinical validation study, comparing the algorithm's output against a reference standard — usually expert radiologist consensus, sometimes confirmed by biopsy, surgery, or longitudinal follow-up — to generate the quantitative performance evidence that anchors the entire submission.
Sensitivity measures the fraction of true-positive cases (e.g., actual pulmonary nodules) the algorithm correctly flags; specificity measures the fraction of true-negative cases it correctly clears. AUC-ROC summarizes the tradeoff between the two across all possible decision thresholds, giving a single number for overall discriminative performance. Positive and negative predictive values (PPV/NPV) further depend on how common the target finding is in the population being screened, which is why study populations must reasonably reflect intended clinical use.
All of these metrics are only as meaningful as the reference standard they are measured against. For findings confirmable by tissue diagnosis (cancer) or a definitive downstream event (stroke), biopsy or clinical outcome anchors the ground truth. For more subjective findings, expert radiologist consensus — often with a panel of 3 or more readers and a defined adjudication process — serves as the reference standard instead.
Many imaging AI submissions use an MRMC design: the same case set is read by a panel of radiologists both with and without AI assistance, allowing the sponsor to show not just standalone algorithm accuracy but the incremental clinical benefit of AI assistance to a human reader — often the more clinically meaningful claim for a device intended to support, rather than replace, a radiologist.
Study sample sizes are driven by statistical power calculations: to detect a clinically meaningful difference in sensitivity or AUC with adequate confidence, imaging AI validation studies commonly require several hundred to several thousand cases, especially when the target finding is relatively rare in the population and multiple subgroups must each be adequately powered.
Performance evidence in hand, the sponsor compiles a comprehensive 510(k) submission: device description, a side-by-side predicate comparison, the full performance dataset and statistical analysis, software documentation appropriate to the device's risk level, cybersecurity documentation, and proposed labeling — all before the FDA review clock even starts.
The device description explains the algorithm's architecture, intended use statement, and indicated patient population and imaging modality. The predicate comparison table lines up intended use and technological characteristics against the chosen predicate, point by point. The performance data section presents the validation study design, dataset characteristics, subgroup breakdowns, and statistical results. Software documentation — scaled to "Basic" or "Enhanced" depending on the device's risk level — details the software development lifecycle, verification and validation testing, and version history. Labeling drafts the instructions for use, warnings, and limitations that will accompany the cleared device.
Since Section 524B of the Food, Drug & Cosmetic Act took effect in October 2023, FDA can refuse to even accept a "cyber device" submission — which includes essentially all networked AI/ML-based software as a medical device — unless it includes a Software Bill of Materials (SBOM), a plan for identifying and addressing post-market cybersecurity vulnerabilities, and evidence the device is designed to provide reasonable assurance of cybersecurity throughout its lifecycle.
This has made cybersecurity documentation a genuine acceptance-review gate rather than a supplementary appendix, and sponsors now build SBOM generation and vulnerability management planning into their development process well before submission, not after.
As of October 2023, FDA can decline to accept for review any cyber device submission lacking a Software Bill of Materials and a cybersecurity management plan under Section 524B of the FD&C Act — a submission-blocking requirement that did not exist for earlier generations of cleared imaging AI devices.
Once accepted for review, FDA works the submission against a statutory clock — a 90-day goal for standard 510(k) review — during which reviewers may pause the clock with Additional Information requests. A clearance decision lets the device be marketed, but for continuously-learning AI, the regulatory relationship with FDA does not end there.
FDA's Medical Device User Fee Amendments (MDUFA) set a goal of 90 days for a standard 510(k) review decision, measured from acceptance. In practice, most submissions receive at least one Additional Information (AI) request — a formal set of reviewer questions or requests for missing data — which pauses ("stops") the clock until the sponsor responds, so calendar time to decision commonly stretches to roughly 150 days or more even when the underlying review workload fits within the 90-day goal.
The eventual decision is either clearance (the device may be marketed for its stated intended use), a request for additional information that keeps the file open, or, less commonly for a well-prepared 510(k), a not-substantially-equivalent determination that sends the sponsor back to the drawing board on either the science or the pathway itself.
Traditional 510(k) devices are locked at clearance: any modification significant enough to affect safety or effectiveness normally requires a new submission. That is a poor fit for AI models designed to keep learning from new data after deployment. FDA's 2023 guidance on Predetermined Change Control Plans (PCCPs) lets a sponsor pre-specify, at the time of original clearance, the scope of future model changes, the retraining and validation protocol, and the performance boundaries within which the device can be updated — without triggering a brand-new 510(k) for every retrained version.
After clearance, cleared devices remain subject to post-market surveillance obligations, including Medical Device Reporting (MDR) of adverse events and malfunctions, and FDA can request real-world performance data to confirm the device continues to perform as validated once used broadly across sites, scanners, and populations beyond the original study.
FDA's 2023 Predetermined Change Control Plan framework allows sponsors of adaptive AI/ML radiology algorithms to pre-authorize a defined envelope of future model updates — a landmark shift acknowledging that "locked" software regulation does not fit devices designed to keep learning after they reach the market.