image:
Image of the entire slide showing detection of important tissue structures such as glands and cells. Credit: Dr Fayyaz Minhas / University of Warwick
view more
Credit: Dr Fayyaz Minhas / University of Warwick
A new study warns that common deep learning systems trained for cancer pathology may rely on hidden shortcuts rather than genuine biological signals.
Artificial intelligence tools are increasingly being developed to predict cancer biology directly from microscopic images, promising faster diagnosis and cheaper testing. But new research from the University of Warwick shows that natural biomedical engineeringsuggests that many of these systems may be using visual shortcuts rather than true biology, raising concerns that some AI pathology tools may now be too unreliable for real-world patient care.
“This is similar to judging the quality of a restaurant by the line of people waiting to get in. It’s a useful shortcut, but it’s not a direct measure of what’s going on in the kitchen.” Dr. Fayyaz Minhas, Associate Professor and Principal Researcher in the Biomedical Prediction Systems (PRISM) Laboratory at the School of Computer Science at the University of Warwick and lead author of the study, said: “Many AI pathology models do the same thing, relying on correlations between biomarkers or obvious tissue features rather than isolating biomarker-specific signals. And these shortcuts often break down as conditions change.”
To reach this conclusion, researchers analyzed more than 8,000 patient samples across four major cancer types: breast, colorectal, lung, and endometrial cancer, and compared the performance of leading machine learning approaches. Although the model often achieved high headline accuracy, the researchers found that this often came from statistical “shortcuts.”
For example, rather than detecting mutations in the BRAF gene that are associated with cancer, the model might learn that BRAF mutations often occur together with another clinical feature, such as microsatellite instability (MSI). The system then learns to use this combination of cues to predict the BRAF state, rather than learning the causative BRAF signal itself. In other words, accurate cancer prediction only works when these biomarkers co-occur, and is less reliable when they do not.
Kim Branson, senior vice president of artificial intelligence and machine learning at GSK and co-author, said: “We found that predicting BRAF mutations by looking at correlated features like MSI is akin to predicting rain by looking at an umbrella. It works, but it doesn’t mean we understand meteorology. Importantly, if our models can’t demonstrate information gain beyond the simple grade assigned by a pathologist, we’re not advancing the field; we’re simply automating shortcuts. Next Generation Pathology AI The roadmap doesn’t necessarily have to be big models; it’s more rigorous evaluation protocols that force algorithms to stop cheating and learn difficult biology.
When the performance of the AI model was evaluated within stratified patient subgroups, such as only high-grade breast cancer or only MSI-positive tumors, accuracy decreased significantly, revealing that the model relied on shortcut signals that disappeared once confounders were controlled.
For certain predictive tasks, the performance advantage of deep learning over human-derived clinical information was modest. The AI system achieved an accuracy score of just over 80% when predicting biomarkers, compared to about 75% when using tumor grade alone. This score has already been evaluated by a pathologist.
Professor Nasir Rajput, Director of the Center for Tissue Image Analysis (TIA) at the University of Warwick and CEO of Warwick spin-out company Histofi, said: “This study highlights an important point regarding the deployment of AI in healthcare: To have real and lasting impact, the value of clinically relevant AI-based predictions must be determined through rigorous bias-aware evaluation, rather than relying solely on headline accuracy that does not account for confounding effects.”
Machine learning methods continue to prove valuable in research, drug development candidate screening, clinical triage, screening, or complementary decision support. However, the researchers argue that future AI tools will need to move beyond correlation-based learning and adopt approaches that explicitly model biological relationships and causal structure. They also want stronger evaluation criteria, such as subgroup testing and comparison to simple clinical baselines, before considering implementation in routine clinical practice.
Dr. Minhas concludes: “This study is not a condemnation of AI in pathology; it is a wake-up call. Current models, while they may work well in controlled settings, rely on statistical shortcuts rather than true biological understanding. Until more robust evaluation criteria are established, these tools should not be viewed as a replacement for molecular testing, and it is imperative that clinicians and researchers understand their limitations and use them with appropriate caution.”
Co-author Professor Sabine Tejpard, Head of the Department of Gastrointestinal Oncology at the University of Leuven, said: “The clinical relevance of new tools requires evidence-based adjustments based on what is accurate, precise, and actionable for individual patients. Too often oncology is dominated by ‘innovations’ that have limited or no impact on patient care and are driven by what can be delivered or sold rather than rigorous evaluation of what is truly relevant to individual patients and their specific characteristics.”
“Progress often requires imperfect first steps, but we must learn from the past and avoid oversimplification and overreach with inappropriate concepts. Complexity and variability are central challenges, but these are also exactly what these new technologies must learn to embrace.”
end
Note to editor
For more information, please contact us below.
Dr Matt Higgs | Media & Communications Officer (Warwick Press Office)
Email: Matt.Higgs@warwick.ac.uk | Phone: +44(0)7880 175403
About research
The paper “Predicting molecular biomarkers from histological images is rife with confounding factors and biases.” natural biomedical engineering. DOI: 10.1038/s41551-026-01616-8
This large-scale analysis was led by first author Dr Muhammad Dawood, a PhD student at the University of Warwick. He is currently a postdoctoral fellow at the University of Oxford.
why is this important
- Biomarkers can guide treatment decisions. If AI tools confuse correlated signals, patients may receive inappropriate treatment.
- Accurate scores can be misleading. This study shows why deeper validation is essential before clinical implementation.
- AI promises faster and cheaper diagnosis, but premature deployment can undermine trust and lead to costly errors.
- The findings signal a shift toward causal, biologically aware AI models that better reflect how diseases actually work.
About the University of Warwick
Founded in 1965, the University of Warwick is a world-leading institution known for its commitment to era-defining innovation across research and education. A connected ecosystem of staff, students and alumni, the University fosters innovative learning, cross-disciplinary collaboration and bold industry partnerships across its state-of-the-art facilities in the UK and global satellite hubs. Here, vibrant thinkers push boundaries, experiment, and challenge conventional wisdom to create a better world.
journal
natural biomedical engineering
Research method
image analysis
Research theme
human tissue sample
Article title
“There are many confounding factors and biases when predicting molecular biomarkers from histological images.”
Article publication date
March 2, 2026
Conflict of interest statement
M.D. conducted this research while a doctoral student at the University of Warwick, UK. MD receives doctoral student support from GSK Inc. KB is an employee of GSK Inc. NR is a founding director, CEO and CSO of Histofy Ltd. FM owns shares in Histofy Ltd but is not involved in its operations. The authors declare that they have no other competing interests.
Disclaimer: AAAS and EurekAlert! We are not responsible for the accuracy of news releases posted on EurekAlert! Use of Information by Contributing Institutions or via the EurekAlert System.
