To our knowledge, this is the first meta-analysis to evaluate the diagnostic performance of a DL algorithm that uses WSI to detect MSI-H in CRC. For the internal validation dataset, patient-based analysis provided sensitivity of 0.88 and specificity of 0.86, while image-based analysis provided sensitivity of 0.81 and specificity of 0.82. The sensitivity AUC was 0.94 and the specificity AUC was 0.84. In contrast, the external validation dataset showed a high sensitivity of 0.93 and a specificity of 0.71 in patient-based analysis. Image-based analysis of external datasets revealed a sensitivity of 0.80 and a specificity of 0.54. The AUC was 0.92 for patient-based analysis and 0.71 for image-based analysis. These results suggest that while the DL algorithm effectively identifies MSI-H in CRC, its performance differs between internal and external validation datasets. The excellent diagnostic performance of deep learning algorithms can be attributed to the ability to automatically learn complex morphological features associated with MSI-H directly from digital pathology slides that traditional pathologists may overlook with the naked eye.37. The higher specificity of the internal validation data set may be the result of consistent data preprocessing, uniform staining, and standardized image acquisition. This helps MSI-H to accurately distinguish it from non-MSI-H cases. In contrast, external validation datasets often result in greater variation due to differences in staining protocols, slide preparation, and image quality, leading to domain shifts and reduced specificity38. These findings highlight the need for standardized data pipelines and the inclusion of multicenter datasets to enhance generalization. Although DL shows an important possibility of MSI-H detection, caution is required due to standardized external validation protocols that may introduce bias due to dataset-specific factors and the absence of standardized external validation protocols. Future research should focus on collaborative frameworks while developing robust and diverse training data sets while adopting cross-validation strategies to reduce overfitting and improve clinical applicability39.
With regard to internal and external validation datasets, it was revealed that the patient-based approach showed higher sensitivity compared to image-based analysis (0.88 vs 0.83, 0.92 vs 0.80). In a patient-based method, each patient is represented as an independent sample with a single WSI image, while an image-based method may include multiple slices from the same patient. With independent sampling, the model captures wider variability and improves predictive performance across diverse patient populations40. The patient-based approach reflects an increased diversity covering variability in tumor type, stage, and treatment response. This diversity improves model generalization by allowing them to learn a wider range of functions, such as tumor staging and demographic characteristics.41. In contrast, image-based training risks overfitting specific features within individual patients. This may limit the applicability of the model to external datasets.42.
In internal and external validation of the AI algorithm, meta-regression analysis revealed no significant statistical differences in sensitivity or specificity between patient-based CNN and non-CNN groups. For example, for non-CNN models, Niehues's study demonstrated that self-monitoring attention-based multi-instance learning models effectively focus on relevant organizational areas.twenty four. Visualization of attentional mechanisms revealed that for MSI prediction, the model focuses minimally on fibromuscular and non-tumor epithelium, while focusing primarily on tumor tissue. However, attentional variance was observed, potentially contributing to the finding that the high attention model did not outweigh the sensitivity or specificity standalone CNN algorithm. Future comparisons of diagnostic performance between different deep learning algorithms are a promising field of exploration.
Note that in the patient-based external validation dataset, the larger tiles (512*512) showed higher specificity compared to the smaller tiles (224*224 or 256*256) (0.91 vs. 0.58). p<0.001). The larger the tile, the better the model's ability to capture local features. This is important for identifying subtle pathological changes. Conversely, small tiles can provide wider contextual information, but may overlook important details43,44. Although the DL algorithm provides promises to improve pathological diagnosis, further research is needed to investigate the impact of tile size on model performance and ensure reliability of clinical applications.
Furthermore, meta-regression analysis using reference standards showed that patient-based internal and external validation was significantly higher in the non-only PCR group than in the only PCR group. However, current evidence shows that PCR exhibits greater diagnostic performance than IHC, particularly in terms of sensitivity and specificity, as a reference standard for identifying MSIs in CRCs.45,46. PCR has higher specificity than IHC for detecting MSI-H, so using IHC as the gold standard will result in a higher false positive rate (i.e., when considered positive by IHCs that are not truly positive). In this situation, as long as the deep learning model detects morphological features associated with IHC positivity in the image, these cases count as “true positives,” and thus overestimates the sensitivity of the model. In contrast, when PCR is used as a reference standard, the model is required to accurately identify PCR-positive cases. This can reduce sensitivity, but it more accurately reflects the actual biological state. Nevertheless, heterogeneity between the study and the relatively few articles in only PCR group may contribute to the potential instability of the results. Therefore, future studies with larger sample sizes are essential to assess the diagnostic performance of different reference standards and achieve more robust findings.
Meanwhile, a previous systematic review by Davri et al.47 and Guitton et al.48has provided valuable insights into the use of DL for CRC diagnosis and prediction of MSI from WSIS. Our study strengthens this foundation by incorporating a broader internal and external datasets for systematic statistical analysis. This approach improves the assessment of model adaptability across different populations. Furthermore, concerns not addressed thoroughly in the existing literature highlight the need to standardize algorithms to mitigate potential overfitting problems during external verification.
Compared to previous meta-analysis by Ying et al. Alam et al. ,Our meta-analysis is the first to predict MSI-H in CRC using WSI. Our study also includes larger sample sizes, and incorporates more research. Ying et al. Meta-analysis uses complex confounding models that combine traditional machine learning, clinical and genomic features, with limited scalability49. Alam et al. The study evaluated MSI predictions in multiple cancer types, including colorectal, stomach, ovarian, and endometrial cancer, but did not perform a pooled analysis of the diagnostic performance of MSI-H-specific DL in CRCs.50. In another meta-analysis, Wang et al. We evaluated AI-based radioactivity of MSI predictions in CRC, but included fewer studies (14 studies) and a limited external validation data set (4 data sets).51The reported AUC was 0.83 and the sensitivity was 0.76, both lower than the AUC of 0.90 and 0.91 sensitivity. Furthermore, Wang et al. Nine of the 12 studies in the analysis relied on PET/CT. This is different from the AI's goal of expensive and cost-effective diagnosis. In contrast, our study shows that WSI-based AI models can efficiently identify MSI-H in CRCs, providing new evidence of clinical applicability and advantages in CRC diagnosis.
High heterogeneity between included studies may have influenced the pooled sensitivity and specificity of DLs in both internal and external validation datasets. Reference standards as sources of heterogeneity in multiple meta-regression identification centers, AI algorithms, analysis methods, magnification, tile size, and reference standards. When it comes to external validation sensitivity, analysis methods, magnification, and tile size were key contributors. Specificity, center, AI algorithm, tile size, and reference standards influenced internal validation, but expansion was the only factor in external validation. However, this heterogeneity can be attributed to other potential factors such as clinical staging of colorectal cancer, dataset size, local population size, WSI image quality, and specimen origin (surgical resection or endoscopic biopsy).
Our results show that the DL-based method achieves high diagnostic performance of MSI-H detection in colorectal cancer in both internal and external datasets. AI can reduce clinician workload, minimize diagnostic errors, and prevent harmful consequences associated with misdiagnosis. However, only one study in the analysis compared AI directly to human performance. Cather etc. For pathologists, we reported sensitivity and specificity of 0.518. Future research should focus on comparative evaluations of AI and human performance, particularly pathologist performance. Beyond diagnostic performance, cost-effectiveness is important for integrating AI models into everyday practice. A virtual metastatic CRC population can save approximately $400 million by combining sensitive AI with confirmatory MSI tests52. The AI model also facilitates treatment initiation, reduces the average time to less than one day, and improves patient outcomes. Once trained, AI systems should minimize maintenance costs while providing valuable insights that can reduce unnecessary treatments and accelerate diagnosis.52. Despite these promising possibilities, some challenges remain. AI models require large, diverse datasets for robust validation and effective integration into everyday clinical workflows. Training these models takes time, often requires hundreds or thousands of annotated images and may include extensive manual labels. Furthermore, concerns about data privacy, model interpretability, and regulatory approval further complicate implementation. Addressing these challenges is essential to ensuring successful and safe adoption of AI in clinical practice50.
Some limitations of this meta-analysis should be considered carefully when interpreting the results. First, the training and validation cohorts of all included models are retrospective and may introduce potential biases. Prospective studies are needed to verify these findings and ensure their applicability in clinical practice.twenty three. Second, some studies used a combination of PCR and IHC as a reference standard. Weak staining of IHC may have missed cases and may be biased towards diagnostic performance to identify MSI-H in CRC53. Third, model training relies heavily on specific open data sets (e.g. TCGA, Quasar, DACHS, etc.), and limited use of local clinical WSIS images for training and validation. This dependency can lead to bias and hinder the assessment of the generalizability of the model. Fourth, we recognize that choosing only the best performance algorithms from multimodel studies may introduce a positive performance bias, as they do not represent the full range of the algorithms tested. To minimize patient overlap among the included studies, we chose to extract only the best-performing algorithms from each study. This can result in performance overestimation. Furthermore, due to limited data availability, we used the estimated maximum Youden index. This could also contribute to biasing performance estimates. It is also important to highlight that Quadas-2 assessments have shown that the “unclear” risk of bias in patient selection for 17 of 19 studies and analytical domains indicates potential spectral bias and selective reporting.
In conclusion, this meta-analysis confirms that the DL algorithm performs better in detecting microsatellite MSI-H in CRC using WSI. However, the low specificity of external validation suggests overfitting and highlights the need for algorithmic standardization to improve generalizability and clinical utility.
