AI shows high accuracy in CT, MRI protocols

Machine Learning


A systematic review and meta-analysis of 23 studies involving approximately 1.2 million imaging orders found that artificial intelligence systems can assign computed tomography and magnetic resonance imaging protocols with an overall accuracy of approximately 85% and perform similarly across several models.

Across analyses, protocol accuracy was 83% for traditional machine learning models, 87% for transformer-based models built on bidirectional encoder representations from transformers, and 86% for large-scale language models. Differences between approaches were not statistically significant.

For research published in American Journal of RoentgenologyResearchers led by Dr. Ethan Sakolansky from the University of Saskatchewan in Canada reviewed studies published between 2017 and 2025 that evaluated artificial intelligence tools designed to automate the assignment of computed tomography or magnetic resonance imaging protocols using information from image requests.

The researchers searched multiple databases through July 2025 and included English-language studies that reported quantitative performance metrics such as accuracy, precision, recall, and F1 score.

Across 23 studies, researchers analyzed 57 performance results representing 30 different models. The training dataset ranged from 1,235 to 559,305 cases, with a median of approximately 61,600 cases. The test dataset ranged from 100 to almost 140,000.

Six studies evaluated applications for computed tomography, 10 evaluated applications for magnetic resonance imaging, and seven included both imaging modalities.

Traditional machine learning models (such as random forests, support vector machines, gradient boosting machines, and deep neural networks) were evaluated in 16 studies and achieved accuracies ranging from 65% to 95%.

Eight studies evaluated transformer-based models built on bidirectional encoder representations from transformers. All include task-specific tweaks using image processing request text. These models achieved accuracy between 78% and 93%, with the highest performing individual model, BioBERT, reaching 93%. Of the 10 best-performing models identified across studies, 6 were BERT-based.

Large-scale language models were evaluated in five studies, including GPT-3.5 Turbo, GPT-4, GPT-4o, and an inference-focused variant. Accuracy ranged from 78% to 92%, with a pooled accuracy of 86%. Only one study used task-specific fine-tuning, which the researchers suggest may partially explain why the trans-based BERT model performed slightly better overall.

Subgroup analyzes showed similar performance across imaging modalities. Accuracy was 83% for computed tomography requests and 85% for magnetic resonance imaging requests. Additionally, accuracy was slightly higher for English requests than for non-English requests.

Investigators have identified several common causes of protocol errors. Ambiguous or incomplete request statements, such as ambiguous symptoms or lack of anatomical details, often led to incorrect protocols being selected. Data imbalance also played a role, as models trained on datasets dominated by common protocols performed poorly when predicting rare protocol categories.

Furthermore, some artificial intelligence “errors” reflect clinically acceptable alternative protocols rather than obvious mistakes, highlighting variability in radiologist decision-making.

Despite these limitations, current performance levels suggest that if deployed carefully, artificial intelligence can help streamline radiology workflows, the authors said. A hybrid system in which an algorithm automatically assigns a routine protocol and uncertain cases are referred to a radiologist for review may provide the most practical short-term approach.

“This tool shows strong potential to help streamline radiologists’ workflows, perhaps through a hybrid AI-radiology approach,” the researchers wrote.

The authors said future research should focus on prospective clinical trials and further development of fine-tuned large-scale language models.

sauce: American Journal of Roentgenology



Source link