Introduction to using symbolic regression for interpretable machine learning in healthcare

Machine Learning


  • Dong, J. & Zhong, J. Recent advances in symbolic regression. ACM Comput. Surv. 57, 37 (2025). Dong & Zhong provide a thorough review of foundational and contemporary symbolic regression techniques, discussing their history, how they function, their strengths and weaknesses, relevant software tools, and examples of where they see use.

    Article 

    Google Scholar 

  • Angelis, D., Sofos, F. & Karakasidis, T. E. Artificial intelligence in physical sciences: symbolic regression trends and perspectives. Arch. Comput. Methods Eng. 30, 3845–3865 (2023).

    Article 

    Google Scholar 

  • Aldeia, G. S. I. & de França, F. O. Interpretability in symbolic regression: a benchmark of explanatory methods using the Feynman data set. Genet. Program. Evolable Mach. 23, 309–349 (2022).

    Article 

    Google Scholar 

  • Ennab, M. & Mcheick, H. Enhancing interpretability and accuracy of AI models in healthcare: a comprehensive review on challenges and future directions. Front. Robot. AI 11, 1444763 (2024).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar 

  • Lemos, P., Jeffrey, N., Cranmer, M., Ho, S. & Battaglia, P. Rediscovering orbital mechanics with machine learning. Mach. Learn. Sci. Technol. 4, 045002 (2023).

    Article 

    Google Scholar 

  • Morales-Alvarado, M., Conde, D., Bendavid, J., Sanz, V. & Ubiali, M. Symbolic regression for precision LHC physics. Preprint at https://doi.org/10.48550/arXiv.2412.07839 (2024).

  • Shojaee, P. et al. LLM-SRBench: a new benchmark for scientific equation discovery with large language models. Preprint at https://doi.org/10.48550/arXiv.2504.10415 (2025).

  • LaCava, W. et al. Contemporary symbolic regression methods and their relative performance. In Proc. Neural Information Processing Systems (NeurIPS) Track on Datasets and Benchmarks Vol 1 (NeurIPS, 2021).

  • Sotiropoulos, D. N., Koronakos, G. & Solanakis, S. V. Evolving transparent credit risk models: a symbolic regression approach using genetic programming. Electronics 13, 4324 (2024).

    Article 

    Google Scholar 

  • Li, Q. et al. Advancing symbolic regression for earth science with a focus on evapotranspiration modeling. NPJ Clim. Atmos. Sci. 7, 321 (2024).

    Article 

    Google Scholar 

  • Abdellaoui, I. A. & Mehrkanoon, S. Symbolic regression for scientific discovery: an application to wind speed forecasting. In Proc. IEEE Symposium Series on Computational Intelligence (SSCI) 01–08 https://doi.org/10.1109/SSCI50451.2021.9659860 (2021).

  • Franca, F. O. et al. Interpretable symbolic regression for data science: analysis of the 2022 competition. Preprint at https://doi.org/10.48550/arXiv.2304.01117 (2023).

  • Cranmer, M. Interpretable machine learning for science with PySR and SymbolicRegression.jl. Preprint at https://doi.org/10.48550/arXiv.2305.01582 (2023).

  • Wang, L., Wu, Z., Sun, K., Li, Z. & Cheng, R. EvoGP: a GPU-accelerated framework for tree-based genetic programming. Preprint at https://doi.org/10.48550/arXiv.2501.17168 (2025).

  • Petersen, B. K. et al. Deep symbolic regression: recovering mathematical expressions from data via risk-seeking policy gradients. Preprint at https://doi.org/10.48550/arXiv.1912.04871 (2021).

  • Tian, Y. et al. Interactive symbolic regression with co-design mechanism through offline reinforcement learning. Nat. Commun. 16, 3930 (2025).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar 

  • Shojaee, P., Meidani, K., Barati Farimani, A. & Reddy, C. Transformer-based planning for symbolic regression. Adv. Neural Inf. Process. Syst. 36, 45907–45919 (2023).

    Article 

    Google Scholar 

  • Alaa, A. M. & van der Schaar, M. Demystifying black-box models with symbolic metamodels. In Proc. Advances in Neural Information Processing Systems (eds Wallach, H. et al.) Vol. 32 (Curran Associates, Inc., 2019).

  • Jin, Y., Fu, W., Kang, J., Guo, J. & Guo, J. Bayesian symbolic regression. Preprint at https://doi.org/10.48550/arXiv.1910.08892 (2020).

  • Landajuela, M. et al. A unified framework for deep symbolic regression. Adv. Neural Inf. Process. Syst. 35, 33985–33998 (2022).

    Article 

    Google Scholar 

  • Albano, G., Giorno, V., Román-Román, P. & Torres-Ruiz, F. Study of a general growth model. Preprint at https://arxiv.org/abs/2402.00882v1 (2024).

  • Goethals, S., Martens, D. & Evgeniou, T. Manipulation risks in explainable AI: the implications of the disagreement problem. In Proc. Machine Learning and Principles and Practice of Knowledge Discovery in Databases (eds Meo, R. & Silvestri, F.) 185–200 https://doi.org/10.1007/978-3-031-74633-8_12 (Springer Nature, 2025).

  • Krishna, S. et al. The disagreement problem in explainable machine learning: a practitioner’s perspective. Preprint at https://doi.org/10.48550/arXiv.2202.01602 (2024).

  • Breiman, L. Statistical modeling: the two cultures (with comments and a rejoinder by the author). Stat. Sci. 16, 199–231 (2001).

    Article 

    Google Scholar 

  • Liu, J. & Rudin, C. User-Guided Interpretable Models: Rashomon Effect, Interaction, and Computation. Harv. Data Sci. Rev. 7, 3 (2025).

  • Wilstrup, C. & Kasak, J. Symbolic regression outperforms other models for small data sets. Preprint at https://doi.org/10.48550/arXiv.2103.15147 (2021).

  • Eichhorn, S., Mohapatra, A. & Goebel, C. PISR: physics-informed symbolic regression for predicting power system voltage. In Proc. 16th ACM International Conference on Future and Sustainable Energy Systems 92–107 https://doi.org/10.1145/3679240.3734622 (Association for Computing Machinery, 2025).

  • Cho, S., Kim, D. & Hazlett, C. Inference at the Data’s Edge: Gaussian Processes for Estimation and Inference in the Face of Extrapolation Uncertainty. Political Analysis 1–20 https://doi.org/10.1017/pan.2026.10032 (2026).

  • Ziyin, L., Hartwig, T. & Ueda, M. Neural networks fail to learn periodic functions and how to fix it. In Proc. Advances in Neural Information Processing Systems Vol. 33, 1583–1594 (Curran Associates, Inc., 2020).

  • Li, W. et al. A neural-guided dynamic symbolic network for exploring mathematical expressions from data. Preprint at https://doi.org/10.48550/arXiv.2309.13705 (2024).

  • Sahoo, S., Lampert, C. & Martius, G. Learning equations for extrapolation and control. In Proc. 35th International Conference on Machine Learning 4442–4450 (PMLR, 2018).

  • Ferrari, D., Guidetti, V., Wang, Y. & Curcin, V. Multi-objective symbolic regression to generate data-driven, non-fixed structure and intelligible mortality predictors using EHR: binary classification methodology and comparison with state-of-the-art. In AMIA’s Annual Symposium Proceedings Vol. 2022, 442–451 (AMIA, 2023)

  • Russeil, E. et al. Multi-view symbolic regression. Preprint at https://doi.org/10.48550/arXiv.2402.04298 (2024).

  • Hasanzadeh, F. et al. Bias recognition and mitigation strategies in artificial intelligence healthcare applications. NPJ Digit. Med. 8, 154 (2025).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar 

  • Imai Aldeia, G. S. et al. Call for action: towards the next generation of symbolic regression benchmark. in Proc. Genetic and Evolutionary Computation Conference Companion 2529–2538 (Association for Computing Machinery, 2025). Aldeia et al. use modified benchmark problems to provide an objective overview of a wide range of state-of-the-art SR algorithms.

  • Radwan, Y. A., Kronberger, G. & Winkler, S. A comparison of recent algorithms for symbolic regression to genetic programming. In Proc. Computer Aided Systems Theory – EUROCAST 2024 (eds Quesada-Arencibia, A., Affenzeller, M. & Moreno-Díaz, R.) 157–171 https://doi.org/10.1007/978-3-031-82949-9_15 (Springer Nature, 2025).

  • Sun, C., Shen, S., Tao, W., Xue, D. & Zhou, Z. Noise-resilient symbolic regression with dynamic gating reinforcement learning. In Proc. AAAI Conference on Artificial Intelligence Vol. 39, 20690–20698 (AAAI, 2025).

  • Wang, X., Zhao, H., Zhao, Q. & Jiang, B. Deep reinforcement learning-based symbolic regression for PDE discovery using spatio-temporal rewards. In Proc. IEEE 20th International Conference on Automation Science and Engineering (CASE) 3256–3261 https://doi.org/10.1109/CASE59546.2024.10711602 (2024).

  • Cao, X. & Yousefzadeh, R. Extrapolation and AI transparency: why machine learning models should reveal when they make decisions beyond their training. Big Data Soc. 10, 20539517231169731 (2023).

    Article 

    Google Scholar 

  • Epperson, J. F. On the runge example. Am. Math. Mon. 94, 329–341 (1987).

    Article 

    Google Scholar 

  • Kammerer, L., Kronberger, G. & Winkler, S. Bias and variance analysis of contemporary symbolic regression methods. Appl. Sci. 14, 11061 (2024).

    Article 

    Google Scholar 

  • Boulesteix, A.-L. Ten simple rules for reducing overoptimistic reporting in methodological computational research. PLoS Comput. Biol. 11, e1004191 (2015).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar 

  • DeCamp, M. & Tilburt, J. C. Why we cannot trust artificial intelligence in medicine. Lancet Digit. Health 1, e390 (2019).

    Article 
    PubMed 

    Google Scholar 

  • Arbelaez Ossa, L. et al. Re-focusing explainability in medicine. Digit. Health 8, 20552076221074488 (2022).

    PubMed 
    PubMed Central 

    Google Scholar 

  • La Cava, W. G. et al. A flexible symbolic regression method for constructing interpretable clinical prediction models. NPJ Digit. Med. 6, 1–14 (2023). La Cava et al. conduct the first study utilizing SR for EHR prediction modeling and perform a thorough scientific analysis to demonstrate the viability of their approach.

    Article 

    Google Scholar 

  • Arina, P. et al. Mortality prediction after major surgery in a mixed population through machine learning: a multi-objective symbolic regression approach. Anaesthesia 80, 551–560 (2025).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar 

  • Ferrari, D. et al. Using interpretable machine learning to predict bloodstream infection and antimicrobial resistance in patients admitted to ICU: early alert predictors based on EHR data to guide antimicrobial stewardship. PLoS Digit. Health 3, e0000641 (2024).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar 

  • Afzal, W. Search-based Approaches to Software Fault Prediction and Software Testing, Licentiate of Engineering thesis, Mälardalen University (2009).

  • Sorour, S. S., Saleh, C. A. & Shazly, M. Integrating machine learning and symbolic regression for predicting damage initiation in hybrid FRP bolted connections. Sci. Rep. 15, 18564 (2025).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar 

  • de Rooij, M., van Riel, N. A. W. & O’Donovan, S. D. Conditional universal differential equations capture population dynamics and interindividual variation in c-peptide production. NPJ Syst. Biol. Appl. 11, 84 (2025).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar 

  • Hossain, M. A., Elmoselhi, H., Elshorbagy, A. A. & Shoker, A. The Sask formula to estimate glomerular filtration rate in renal transplant patients. Nephron Clin. Pract. 117, c135–c150 (2011).

    Article 
    PubMed 

    Google Scholar 

  • Fong, K. S. & Motani, M. Symbolic regression for discovery of medical equations: a case study on glomerular filtration rate estimation equations. In Proc. IEEE Conference on Artificial Intelligence (CAI) 1373–1379 https://doi.org/10.1109/CAI59869.2024.00245 (2024).

  • Hidalgo, J. I. et al. Modeling glycemia in humans by means of grammatical evolution. Appl. Soft Comput. 20, 40–53 (2014).

    Article 

    Google Scholar 

  • Chiavegatto Filho, A., Batista, A. F. D. M. & dos Santos, H. G. Data leakage in health outcomes prediction with machine learning. Comment on “prediction of incident hypertension within the next year: prospective study using statewide electronic health records and machine learning”. J. Med. Internet Res. 23, e10969 (2021).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar 

  • Han, L. Addressing distribution shift for robust and trustworthy prediction and causal inference in clinical AI settings. JAMA Netw. Open 8, e2513705 (2025).

    Article 
    PubMed 

    Google Scholar 

  • Subasri, V. et al. Detecting and remediating harmful data shifts for the responsible deployment of clinical AI models. JAMA Netw. Open 8, e2513685 (2025).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar 

  • Warraich, H. J., Tazbaz, T. & Califf, R. M. FDA perspective on the regulation of artificial intelligence in health care and biomedicine. JAMA 333, 241–247 (2025).

    Article 
    PubMed 

    Google Scholar 

  • Shah, N. H., Pfeffer, M. A. & Ghassemi, M. The need for continuous evaluation of artificial intelligence prediction algorithms. JAMA Netw. Open 7, e2433009 (2024).

    Article 
    PubMed 

    Google Scholar 



  • Source link