Researchers develop versatile machine learning tools to automate complex clinical diagnoses

Machine Learning


A research team funded by the US National Institutes of Health (NIH) has developed a versatile machine learning model that could one day vastly expand what medical scans can tell us about disease. Scientists used a tool called Merlin to evaluate 3D abdominal computed tomography (CT) scans to accomplish tasks as simple as identifying anatomical features to as complex as predicting the onset of disease years in advance. Despite being developed as a general-purpose CT model, Merlin outperforms similar automation tools at the tasks it was specifically built to handle.

The team trained the model on a unique set of patient CT scans linked to radiology reports and medical diagnosis codes collected from Stanford University School of Medicine. Researchers note that this is the largest collection of abdominal CT data to date.

Such rich datasets are needed to push the boundaries of what artificial intelligence models can achieve in healthcare. This study demonstrates how carefully crafted training data can significantly streamline workflows and enable surprising insights to support clinical decision-making. ”


Dr. Bruce Tromberg, Director, National Institute of Biomedical Imaging and Bioengineering (NIBIB), NIH

CT is a common form of medical imaging and is often performed in the early stages of a medical evaluation. Obtaining a diagnosis requires a radiologist to interpret the results and often also requires additional tests and clinical evaluation. At baseline, this process is time-consuming and will only become more cumbersome given the growing physician shortage in the United States.

“With Merlin, we have the potential to go beyond traditional radiology and jump directly from imaging to possible diagnoses. And that’s just one potential application,” said co-lead author Dr. Louis Blankemeyer, who conducted the study while a graduate student at Stanford University.

Merlin represents a new class of models, commonly referred to as foundational models, that are trained using large unlabeled datasets spanning many types of information.

In the new study, researchers tested Merlin across six broad activity categories across more than 750 individual tasks involving diagnosis, prognosis, and quality assessment.

To make Merlin capable of a wide range of tasks, researchers first trained it on a clinical data set that combined more than 15,000 3D abdominal CT scans and radiology reports and nearly 1 million diagnosis codes. Merlin used this information as learning material to learn about the relationship between visual and textual data.

The researchers then asked Marlin about more than 50,000 never-before-seen abdominal CT scans from one of four different hospitals to see how well the model matched the human-generated conclusions associated with each scan.

“While Merlin tackled some tasks head-on, such as predicting diagnostic codes, other more complex tasks, such as creating radiology reports from scratch or identifying and outlining organs in 3D space, required additional training,” said co-first author Ashwin Kumar, a graduate student at Stanford University.

The team also introduced state-of-the-art models specific to each task type as comparison points.

On average over 692 different diagnosis codes, Merlin was able to predict which of two scans was more likely to be associated with a particular code with over 81% probability, outperforming several variants of the other two models. On a subset of 102 codes, Merlin’s performance increased to 90%.

In another category, the research team urged Merlin to predict the onset of chronic diseases such as diabetes, osteoporosis, and heart disease in healthy patients based solely on CT scans.

The study authors found that when comparing scans from different subjects, Merlin was able to identify patients at high risk of developing a particular disease over the next five years 75% of the time, compared to 68% for other models. These findings suggest that the model can detect important features in scans that may be lost to the human eye, suggesting that this tool could help identify new biomarkers of disease, Blankemeyer explained.

The researchers upped the ante by asking Merlin to interpret a CT scan of his chest. A CT scan of the chest is a completely absent part of the CT research documentation. Merlin’s unique ability to identify generalizable features of disease enabled it to perform as well or better than models trained on chest scans alone.

Despite being a jack-of-all-trades, Merlin outperformed or matched specialist models in every task. The authors attribute Merlin’s magical touch to its architecture and training data, which allowed it to process complex 3D scans and build associations between visual and textual information.

The researchers have high hopes that their approach will be able to leverage precedent and gain regulatory approval for simpler tasks, but they also plan to improve Merlin to better handle more complex challenges such as reporting.

Although this tool is powerful out of the box, we encourage users to use their own data to fine-tune the model to address their specific needs.

“Our model and data provide the community with a robust backbone on which to build,” said lead author Dr. Akshay Chaudhary, professor of radiology and biomedical data science at Stanford University. “The sky is the limit from here.”

This research was supported by NIBIB through grants R01EB002524 and P41EB027060, by the Medical Imaging Data Resource Center (MIDRC) under contract 75N92020C00021, and by the National Health, Lung, and Blood Institute (NHLBI) through grants R01HL167974 and R01HL169345. and was supported by the National Institute of Arthritis. Musculoskeletal and Skin Diseases (NIAMS) grants R01AR077604 and R01AR079431.

sauce:

National Institutes of Health (NIH)

Reference magazines:

Blankemeyer, L. Others. (2026). Merlin: Computed tomography vision – language-based models and datasets. Nature. DOI: 10.1038/s41586-026-10181-8. https://www.nature.com/articles/s41586-026-10181-8



Source link