Assessing the robustness and responsiveness of large-scale frontier models in medical AI applications

Applications of AI


  • Singhal, K. et al. Large-scale language models encode clinical knowledge. nature 620172–180 (2023).

    Article CAS PubMed PubMed Central Google Scholar

  • Gu, Y. et al. Pre-training domain-specific language models for biomedical natural language processing. in ACM Transactions on Computing for Healthcare (HEALTH) (Editors Lee, I. & Stankovic, JA) 31−23 (Association for Computing Machinery, 2022).

  • Nori, H. Sequential diagnosis using language models. Preprint available at https://arxiv.org/abs/2506.22405 (2025).

  • Open AI. Introducing GPT-5. https://openai.com/index/introducing-gpt-5/ (2025).

  • Saab, K. et al. Functions of the Gemini model in medicine. Preprint available at https://arxiv.org/abs/2404.18416 (2024).

  • Tu, T. et al. Towards conversational diagnostic AI. Preprint available at https://arxiv.org/abs/2401.05654 (2024).

  • Wang, S. et al. LINS: A general medical Q&A framework to improve the quality and reliability of LLM-generated responses. nut. common. 169076 (2025).

    Article CAS PubMed PubMed Central Google Scholar

  • Arora, RK et al. HealthBench: Evaluating large-scale language models for improving human health. Preprint available at https://arxiv.org/abs/2505.08775 (2025).

  • Handler, R., Sharma, S. & Hernandez-Boussard, T. The fragile intelligence of GPT-5 in medicine. nut. medicine. 313968–3970 (2025).

    Article CAS PubMed Google Scholar

  • Farquhar, S., Kossen, J., Kuhn, L. & Gal, Y. Hallucination detection in large-scale language models using semantic entropy. nature 630625–630 (2024).

    Article CAS PubMed PubMed Central Google Scholar

  • Jin, Q. et al. Hidden flaws behind the expert-level accuracy of multimodal GPT-4 vision in healthcare. NPJ digit. medicine. 7190 (2024).

    Article PubMed PubMed Central Google Scholar

  • Pfau, J., Merrill, W. & Bowman, S.R. Thinking point by point: Hidden computation in transformer language models. in 1st Conference on Language Modeling (COLM) https://openreview.net/forum?id=NikbrdtYvG (2024).

  • Gayhos, R. et al. Shortcut learning in deep neural networks. nut. Mach. intelligence. 2665–673 (2020).

    Article Google Scholar

  • Acosta, JN, Falcone, GJ, Rajpurkar, P. & Topol, EJ Multimodal biomedical AI. nut. medicine. 281773–1784 (2022).

    Article CAS PubMed Google Scholar

  • Goodfellow, IJ, Shlens, J. & Szegedy, C. Explaining and leveraging adversarial examples. Preprint available at https://arxiv.org/abs/1412.6572 (2015).

  • Szegedi, C. et al. Interesting properties of neural networks. Preprint available at https://arxiv.org/abs/1312.6199 (2013).

  • New England Medical Journal: Image Challenge. https://www.nejm.org/image-challenge (2026).

  • JAMA Network Clinical Issues. https://jamanetwork.com/collections/44038/clinical-challenge (2026).

  • Comanici, G. et al. Gemini 2.5: Pushing the frontiers with advanced reasoning, multimodality, long context, and next-generation agent capabilities. Preprint available at https://arxiv.org/abs/2507.06261 (2025).

  • Human. Claude 3.5 Sonnet. https://www.anthropic.com/news/claude-3-5-sonnet (2024).

  • Open AI. GPT-4o system card. https://openai.com/index/gpt-4o-system-card/ (2024).

  • Open AI. OpenAI o3 and o4-mini system cards. https://openai.com/index/o3-o4-mini-system-card/ (2025).

  • Wei, J. et al. Thought chain prompts draw inferences in large-scale language models. in NIPS’22: Proceedings of the 36th International Conference on Neural Information Processing Systems 24824−24837 (Koyejo, S. et al., eds.) (Curran Associates, 2022).

  • Lau, JJ, Gayen, S., Ben Abacha, A., Demner-Fushman, D. A dataset of clinically generated visual questions and answers about radiology images. Science. data 5180251 (2018).

    Article PubMed PubMed Central Google Scholar

  • Hu, Y. et al. OmniMedVQA: A new large-scale comprehensive evaluation benchmark for medical LVLM. in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) https://doi.org/10.1109/CVPR52733.2024.02093 (IEEE, 2024).

  • Johnson, AE et al. MIMIC-CXR is an anonymized public chest radiograph database with free text reports. Science. data 6317 (2019).

    Article PubMed PubMed Central Google Scholar

  • He, X., Zhang, Y., Mou, L., Xing, E., Xie, P. PathVQA: 30,000+ Questions for Medical Visual Question Answering. Preprint available at https://arxiv.org/abs/2003.10286 (2020).

  • Liu, B. et al. SLAKE: A semantically labeled knowledge enrichment dataset for medical visual question answering. Preprint at https://arxiv.org/abs/2102.09542 (2021).

  • Zhang, X. et al. PMC-VQA: Visual instruction tuning for medical visual question answering. Preprint available at https://arxiv.org/abs/2305.10415 (2023).

  • Yue, X. et al. MMMU: A large-scale, multidisciplinary, multimodal understanding and reasoning benchmark for expert AGI. in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) https://doi.org/10.1109/CVPR52733.2024.00913 (IEEE, 2024).

  • Fleiss, JL Measuring the agreement of nominal scales between many raters. Psychol. Bull. 76378–382 (1971).

    Article Google Scholar

  • Wu, Z. et al. DeepSeek-VL2: An expert mixed visual language model for advanced multimodal understanding. Preprint available at https://arxiv.org/abs/2412.10302 (2024).

  • Bai, S. et al. Qwen3-VL Technical Report. Preprint available at https://arxiv.org/abs/2511.21631 (2025).

  • Lee, C. et al. LLaVA-Med: Training large-scale language and visual assistants for biomedicine in one day. in NIPS ʼ23: Proceedings of the 37th International Conference on Neural Information Processing Systems (Oh, A. et al. eds.) 28541−28564 (Curran Associates, 2023).

  • Sellergren, A. et al. MedGemma Technical Report. Preprint available at https://arxiv.org/abs/2507.05201 (2025).



  • Source link