Foundation models become a new trend in medical technology, but what exactly is a "foundation model"?
Over the past year, companies such as GE HealthCare and Philips have successively deployed MRI foundation models, and the FDA has also begun exploring related labeling. However, experts point out that the definition of foundation models is vague, and it is difficult to assess whether existing tools truly benefit radiologists and patients.

Editor's Note: This article is the first of a two-part series on foundation model applications in the medical technology industry. The second part will be published on Tuesday.
An increasing number of medical device companies are promoting "foundation models"—an artificial intelligence technology that can adapt to multiple tasks. However, there are still many questions regarding the application of this technology in the medical technology industry.
Over the past year, GE HealthCare has promoted its MRI research foundation model, Philips announced plans to collaborate with NVIDIA to build an MRI foundation model, and multiple abstracts at last year's Radiological Society of North America (RSNA) annual meeting focused on how to evaluate and improve foundation models. The U.S. Food and Drug Administration (FDA) has also updated its AI medical device database, indicating that it is exploring methods to identify and label devices that incorporate foundation models.
However, experts say that the definition of foundation models is still unclear, and it is difficult to determine whether currently available tools are truly helping radiologists and patients.
What is a foundation model?
Magdalini Paschali, a postdoctoral scholar in the Department of Radiology at Stanford University, points out that foundation models have several key characteristics: they are trained on large-scale datasets, most of which are unlabeled; they can handle multiple data types, such as images, text, medical history, and genomics; and finally, they can handle multiple tasks, for example, a model can detect diseases it has never seen during training.

Earlier this year, Paschali published a paper in RSNA's journal Radiology, attempting to define the technology more clearly.
Akshay Chaudhari, Assistant Professor of Radiology and Biomedical Data Science at Stanford University, says that in practice, "anything can be defined as a foundation model."
Chaudhari says the term "foundation model" was first coined at Stanford University in 2021. In the medical field, one of the earliest versions was Med-PaLM, launched by Google at the end of 2022, a large language model designed to answer medical questions.
Chaudhari says foundation models began to become more prominent at RSNA in 2023.
Traditional deep learning models used in radiology (such as those used to detect pneumonia) focus on specific health conditions and rely on labeled data. For example, radiologists will review images, circle instances of pneumonia, or mark them in text reports, explains Nina Kottler, Associate Chief Medical Officer of Clinical AI at Radiology Partners.
Chaudhari points out that foundation models are trained on millions of images, not thousands, and requiring labeled data is simply unrealistic.
Are foundation models more accurate?
Some medical device developers claim that foundation models are more accurate than narrow AI models. For example, Aidoc, which provides triage software for radiology, says that foundation models enable faster development of more accurate AI tools.
Experts say the accuracy of a device depends on how it is built. Stanford's Paschali says that a foundation model, if used "out of the box" without any specialization or additional training, may perform worse than a specific AI tool, but once it sees some examples and context, it may perform better.
"The best we can do is look at the summary statements in the FDA clearance documents. At least in the market, we haven't really seen the benefits of these foundation models yet."

Akshay Chaudhari
Assistant Professor of Radiology and Biomedical Data Science at Stanford University
Kottler says that because foundation models are trained on large amounts of data, they may perform better at detecting rare events, such as brain aneurysms.
"When only a few people have a certain disease, finding it is like looking for a needle in a haystack," Kottler says. "You need a very accurate model to do that."

Another advantage of foundation models is the ability to develop other AI models faster. Kottler says that building a traditional narrow AI model might take six months to clean data, label data, and train, while different iterations of a foundation model can be completed within weeks.
However, in practice, it is difficult to know whether these advantages have translated into actual benefits for patients or their care teams. Stanford's Chaudhari says that for foundation models that have received FDA clearance, there is almost no public information supporting companies' claims.
"The best we can do is look at the summary statements in the FDA clearance documents," he says. "At least in the market, we haven't really seen the benefits of these foundation models yet."
Evaluating foundation models
Currently, FDA-cleared foundation models are designed to address specific tasks, such as Aidoc's rib fracture triage tool, which is built on the company's foundation model. Aidoc obtained 510(k) clearance for its model using an older rib fracture triage tool as a predicate device. The 510(k) process requires manufacturers to demonstrate that their device is substantially equivalent to a predicate device that can be legally sold in the United States.
Chaudhari says that for broader models that integrate language and images or video, there are currently no guidelines.
Some hospitals have proposed systems for evaluating AI models, but these systems are not perfect. For example, hospitals will identify a need, such as a model to identify pneumonia in X-ray images or a model to draft X-ray reports.
"Then they will collect 1,000 images with known labels from their own site and hold a competition to see which vendor performs best on that dataset," Chaudhari says. "It's crude, but it really is the best method available right now because it's the only way to assess whether local performance is truly suitable for a specific task."

For large academic medical centers with data science teams, this may be feasible, but "many hospitals will simply deploy the model and then get data of questionable quality from it," Chaudhari adds.
Paschali says that to test the accuracy of a foundation model, one must first define metrics and tasks based on what the model claims to do. Hospitals should also test how the model performs across different patient subgroups and different scanner types. Finally, hospitals should "stress test" the model to identify potential issues that may arise, such as with very rare diseases.
"That's why it's very important to work closely with radiologists, because they can help us design comprehensive stress tests," Paschali says. "Because they have seen so many cases, they know which tasks are difficult even for them."
There is hope that, once foundation models are trained on larger datasets covering different states, different types of hospitals, and imaging equipment, the evaluation required before deployment may be less.
"I don't think we've reached that stage yet, at least there's no evidence in the world," Chaudhari says. "But that's exactly the allure these foundation models might bring."
Another goal is to free up radiologists' time to address the ongoing shortage of radiologists in the United States and the growing number of images.
"If we ask them to verify outputs and check everything," Chaudhari says, "does that really fulfill the promise of these models?"