Singhal, K. et al. Large-scale language models encode clinical knowledge. nature 620172–180 (2023).
Gu, Y. et al. Pre-training domain-specific language models for biomedical natural language processing. in ACM Transactions on Computing for Healthcare (HEALTH) (Editors Lee, I. & Stankovic, JA) 31−23 (Association for Computing Machinery, 2022).
Nori, H. Sequential diagnosis using language models. Preprint available at https://arxiv.org/abs/2506.22405 (2025).
Open AI. Introducing GPT-5. https://openai.com/index/introducing-gpt-5/ (2025).
Saab, K. et al. Functions of the Gemini model in medicine. Preprint available at https://arxiv.org/abs/2404.18416 (2024).
Tu, T. et al. Towards conversational diagnostic AI. Preprint available at https://arxiv.org/abs/2401.05654 (2024).
Wang, S. et al. LINS: A general medical Q&A framework to improve the quality and reliability of LLM-generated responses. nut. common. 169076 (2025).
Arora, RK et al. HealthBench: Evaluating large-scale language models for improving human health. Preprint available at https://arxiv.org/abs/2505.08775 (2025).
Handler, R., Sharma, S. & Hernandez-Boussard, T. The fragile intelligence of GPT-5 in medicine. nut. medicine. 313968–3970 (2025).
Farquhar, S., Kossen, J., Kuhn, L. & Gal, Y. Hallucination detection in large-scale language models using semantic entropy. nature 630625–630 (2024).
Jin, Q. et al. Hidden flaws behind the expert-level accuracy of multimodal GPT-4 vision in healthcare. NPJ digit. medicine. 7190 (2024).
Pfau, J., Merrill, W. & Bowman, S.R. Thinking point by point: Hidden computation in transformer language models. in 1st Conference on Language Modeling (COLM) https://openreview.net/forum?id=NikbrdtYvG (2024).
Gayhos, R. et al. Shortcut learning in deep neural networks. nut. Mach. intelligence. 2665–673 (2020).
Acosta, JN, Falcone, GJ, Rajpurkar, P. & Topol, EJ Multimodal biomedical AI. nut. medicine. 281773–1784 (2022).
Goodfellow, IJ, Shlens, J. & Szegedy, C. Explaining and leveraging adversarial examples. Preprint available at https://arxiv.org/abs/1412.6572 (2015).
Szegedi, C. et al. Interesting properties of neural networks. Preprint available at https://arxiv.org/abs/1312.6199 (2013).
New England Medical Journal: Image Challenge. https://www.nejm.org/image-challenge (2026).
JAMA Network Clinical Issues. https://jamanetwork.com/collections/44038/clinical-challenge (2026).
Comanici, G. et al. Gemini 2.5: Pushing the frontiers with advanced reasoning, multimodality, long context, and next-generation agent capabilities. Preprint available at https://arxiv.org/abs/2507.06261 (2025).
Human. Claude 3.5 Sonnet. https://www.anthropic.com/news/claude-3-5-sonnet (2024).
Open AI. GPT-4o system card. https://openai.com/index/gpt-4o-system-card/ (2024).
Open AI. OpenAI o3 and o4-mini system cards. https://openai.com/index/o3-o4-mini-system-card/ (2025).
Wei, J. et al. Thought chain prompts draw inferences in large-scale language models. in NIPS’22: Proceedings of the 36th International Conference on Neural Information Processing Systems 24824−24837 (Koyejo, S. et al., eds.) (Curran Associates, 2022).
Lau, JJ, Gayen, S., Ben Abacha, A., Demner-Fushman, D. A dataset of clinically generated visual questions and answers about radiology images. Science. data 5180251 (2018).
Hu, Y. et al. OmniMedVQA: A new large-scale comprehensive evaluation benchmark for medical LVLM. in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) https://doi.org/10.1109/CVPR52733.2024.02093 (IEEE, 2024).
Johnson, AE et al. MIMIC-CXR is an anonymized public chest radiograph database with free text reports. Science. data 6317 (2019).
He, X., Zhang, Y., Mou, L., Xing, E., Xie, P. PathVQA: 30,000+ Questions for Medical Visual Question Answering. Preprint available at https://arxiv.org/abs/2003.10286 (2020).
Liu, B. et al. SLAKE: A semantically labeled knowledge enrichment dataset for medical visual question answering. Preprint at https://arxiv.org/abs/2102.09542 (2021).
Zhang, X. et al. PMC-VQA: Visual instruction tuning for medical visual question answering. Preprint available at https://arxiv.org/abs/2305.10415 (2023).
Yue, X. et al. MMMU: A large-scale, multidisciplinary, multimodal understanding and reasoning benchmark for expert AGI. in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) https://doi.org/10.1109/CVPR52733.2024.00913 (IEEE, 2024).
Fleiss, JL Measuring the agreement of nominal scales between many raters. Psychol. Bull. 76378–382 (1971).
Wu, Z. et al. DeepSeek-VL2: An expert mixed visual language model for advanced multimodal understanding. Preprint available at https://arxiv.org/abs/2412.10302 (2024).
Bai, S. et al. Qwen3-VL Technical Report. Preprint available at https://arxiv.org/abs/2511.21631 (2025).
Lee, C. et al. LLaVA-Med: Training large-scale language and visual assistants for biomedicine in one day. in NIPS ʼ23: Proceedings of the 37th International Conference on Neural Information Processing Systems (Oh, A. et al. eds.) 28541−28564 (Curran Associates, 2023).
Sellergren, A. et al. MedGemma Technical Report. Preprint available at https://arxiv.org/abs/2507.05201 (2025).
