Training data composition determines machine learning generalization and biological rule discovery

Machine Learning


  • Hummer, A. M., Schneider, C., Chinery, L. & Deane, C. M. Investigating the volume and diversity of data needed for generalizable antibody–antigen ΔΔG prediction. Nat. Comput. Sci. https://doi.org/10.1038/s43588-025-00823-8 (2025).

  • Yang, R., Mao, J. & Chaudhari, P. Does the data induce capacity control in deep learning? In Proc. 39th International Conference on Machine Learning (ICML) (eds Chaudhuri, K. et al.) 25166–25193 (PMLR, 2022).

  • Alvarez-Melis, D. & Fusi, N. Geometric dataset distances via optimal transport. Adv. Neural Inf. Process. Syst. 33, 21428–21439 (2020).

    Google Scholar 

  • Wang, T. & Isola, P. Understanding contrastive representation learning through alignment and uniformity on the hypersphere. In Proc. 37th International Conference on Machine Learning (ICML) (eds Daumé, H. & Singh, A.) 9929–9939 (PMLR, 2020).

  • Quiñonero-Candela, J., Sugiyama, M., Schwaighofer, A. & Lawrence, N. D. Dataset Shift in Machine Learning (The MIT Press, 2008).

  • Li, X. L., Liu, B. & Ng, S. K. Negative training data can be harmful to text classification. In Proc. 2010 Conference on Empirical Methods in Natural Language Processing (eds Li, H. & Màrquez, L.) 218–228 (Association for Computational Linguistics, 2010).

  • Wunsch, H., Kübler, S. & Cantrell, R. Instance sampling methods for pronoun resolution. In Proc. International Conference on Recent Advances in Natural Language Processing (RANLP-2009) (eds Angelova, G. & Mitkov, R.) 478–483 (Association for Computational Linguistics, 2009).

  • Saeidi, M., Kulkarni, R., Togia, T. & Sama, M. The effect of negative sampling strategy on capturing semantic similarity in document embeddings. In Proc. 2nd Workshop on Semantic Deep Learning (SemDeep-2) (eds Gromann, D. et al.) 1–8 (Association for Computational Linguistics, 2017).

  • Ben-Hur, A. & Noble, W. S. Choosing negative examples for the prediction of protein-protein interactions. BMC Bioinf. 7, S2 (2006).

    Google Scholar 

  • Eid, F.-E., ElHefnawi, M. & Heath, L. S. DeNovo: virus-host sequence-based protein-protein interaction prediction. Bioinformatics 32, 1144–1150 (2016).

    Google Scholar 

  • Tsukiyama, S., Hasan, M. M., Fujii, S. & Kurata, H. LSTM-PHV: prediction of human-virus protein-protein interactions by LSTM with word2vec. Brief. Bioinform. 22, bbab228 (2021).

    Google Scholar 

  • Dens, C., Laukens, K., Bittremieux, W. & Meysman, P. The pitfalls of negative data bias for the T-cell epitope specificity challenge. Nat. Mach. Intell. https://doi.org/10.1038/s42256-023-00727-0 (2023).

    Google Scholar 

  • Sidorczuk, K. et al. Benchmarks in antimicrobial peptide prediction are biased due to the selection of negative data. Brief. Bioinform. 23, bbac343 (2022).

    Google Scholar 

  • Xu, L. et al. Negative sampling for contrastive representation learning: a review. Preprint at https://arxiv.org/abs/2206.00212 (2022).

  • Werner de Vargas, V., Schneider Aranda, J. A., Dos Santos Costa, R., da Silva Pereira, P. R. & Victória Barbosa, J. L. Imbalanced data preprocessing techniques for machine learning: a systematic mapping study. Knowl. Inf. Syst. 65, 31–57 (2023).

    Google Scholar 

  • Loffredo, E., Pastore, M., Cocco, S. & Monasson, R. Restoring balance: principled under/oversampling of data for optimal classification. Preprint at https://arxiv.org/abs/2405.09535 (2024).

  • Akbar, R. et al. Progress and challenges for the machine learning-based design of fit-for-purpose monoclonal antibodies. MAbs 14, 2008790 (2022).

    Google Scholar 

  • Greiff, V., Yaari, G. & Cowell, L. G. Mining adaptive immune receptor repertoires for biological and clinical information using machine learning. Curr. Opin. Syst. Biol. 24, 109–119 (2020).

    Google Scholar 

  • Wilman, W. et al. Machine-designed biotherapeutics: opportunities, feasibility and advantages of deep learning in computational antibody discovery. Brief. Bioinform. 23, bbac267 (2022).

    Google Scholar 

  • Khetan, R. et al. Current advances in biopharmaceutical informatics: guidelines, impact and challenges in the computational developability assessment of antibody therapeutics. MAbs 14, 2020082 (2022).

    Google Scholar 

  • Fernández-Quintero, M. L. et al. Assessing developability early in the discovery process for novel biologics. MAbs 15, 2171248 (2023).

    Google Scholar 

  • Chinery, L. et al. Baselining the buzz trastuzumab-HER2 affinity, and beyond. Preprint at bioRxiv https://doi.org/10.1101/2024.03.26.586756 (2024).

  • Ehling, R. A. et al. Synthetic coevolution reveals adaptive mutational trajectories of neutralizing antibodies and SARS-CoV-2. Preprint at bioRxiv https://doi.org/10.1101/2024.03.28.587189 (2024).

  • O’Donnell, T. J. et al. Reading the repertoire: Progress in adaptive immune receptor analysis using machine learning. Cell Syst 15, 1168–1189 (2024).

    Google Scholar 

  • Schneider, C., Buchanan, A., Taddese, B. & Deane, C. M. DLAB: deep learning methods for structure-based virtual screening of antibodies. Bioinformatics 38, 377–383 (2022).

    Google Scholar 

  • Krützfeldt, L.-M., Schubach, M. & Kircher, M. The impact of different negative training data on regulatory sequence predictions. PLoS ONE 15, e0237412 (2020).

    Google Scholar 

  • Robert, P. A. et al. Unconstrained generation of synthetic antibody-antigen structures to guide machine learning methodology for antibody specificity prediction. Nat. Comput. Sci. 2, 845–865 (2022).

    Google Scholar 

  • Montemurro, A., Jessen, L. E. & Nielsen, M. NetTCR-2.1: lessons and guidance on how to develop models for TCR specificity predictions. Front. Immunol. 13, 1055151 (2022).

    Google Scholar 

  • Grazioli, F. et al. On TCR binding predictors failing to generalize to unseen peptides. Front. Immunol. 13, 1014256 (2022).

    Google Scholar 

  • Deng, L. et al. Performance comparison of TCR-pMHC prediction tools reveals a strong data dependency. Front. Immunol. 14, 1128326 (2023).

    Google Scholar 

  • Gao, Y., Gao, Y., Dong, K., Wu, S. & Liu, Q. Reply to: the pitfalls of negative data bias for the T-cell epitope specificity challenge. Nat. Mach. Intell. 5, 1063–1065 (2023).

    Google Scholar 

  • Xu, J. L. & Davis, M. M. Diversity in the CDR3 region of V(H) is sufficient for most antibody specificities. Immunity 13, 37–45 (2000).

    Google Scholar 

  • Davis, M. M. & Bjorkman, P. J. T-cell antigen receptor genes and T-cell recognition. Nature 334, 395–402 (1988).

    Google Scholar 

  • Akbar, R. et al. A compact vocabulary of paratope-epitope interactions enables predictability of antibody-antigen binding. Cell Rep. 34, 108856 (2021).

    Google Scholar 

  • Dunbar, J. et al. SAbDab: the structural antibody database. Nucleic Acids Res. 42, D1140–D1146 (2014).

    Google Scholar 

  • Shugay, M. et al. VDJdb: a curated database of T-cell receptor sequences with known antigen specificity. Nucleic Acids Res. 46, D419–D427 (2018).

    Google Scholar 

  • Tickotsky, N., Sagiv, T., Prilusky, J., Shifrut, E. & Friedman, N. McPAS-TCR: a manually curated catalogue of pathology-associated T cell receptor sequences. Bioinformatics 33, 2924–2929 (2017).

    Google Scholar 

  • Swindells, M. B. et al. abYsis: integrated antibody sequence and structure-management, analysis, and prediction. J. Mol. Biol. 429, 356–364 (2017).

    Google Scholar 

  • Raybould, M. I. J., Kovaltsuk, A., Marks, C. & Deane, C. M. CoV-AbDab: the coronavirus antibody database. Bioinformatics 37, 734–735 (2021).

    Google Scholar 

  • Mahajan, S. et al. Epitope specific antibodies and T cell receptors in the immune epitope database. Front. Immunol. 9, 2688 (2018).

    Google Scholar 

  • Chen, V. et al. Best practices for interpretable machine learning in computational biology. Preprint at bioRxiv https://doi.org/10.1101/2022.10.28.513978 (2022).

  • Sandve, G. K. & Greiff, V. Access to ground truth at unconstrained size makes simulated data as indispensable as experimental data for bioinformatics methods development and benchmarking. Bioinformatics 38, 4994–4996 (2022).

    Google Scholar 

  • Khan, A. et al. Toward real-world automated antibody design with combinatorial Bayesian optimization.Cell Rep. Methods 3, 100374 (2023).

    Google Scholar 

  • Zhang, C., Bzikadze, A. V., Safonova, Y. & Mirarab, S. A scalable model for simulating multi-round antibody evolution and benchmarking of clonal tree reconstruction methods. Front. Immunol. 13, 1014439 (2022).

    Google Scholar 

  • Akbar, R. et al. In silico proof of principle of machine learning-based antibody design at unconstrained scale. MAbs 14, 2031482 (2022).

    Google Scholar 

  • Sundararajan, M., Taly, A. & Yan, Q. Gradients of counterfactuals. Preprint at https://arxiv.org/abs/1611.02639 (2016).

  • Karim, M. R. et al. Explainable AI for bioinformatics: methods, tools and applications. Brief. Bioinform. 24, bbad236 (2023).

    Google Scholar 

  • Shrikumar, A., Greenside, P. & Kundaje, A. Learning important features through propagating activation differences. In Proc. 34th International Conference on Machine Learning (eds Precup, D. & Teh, Y. W.) 3145–3153 (PMLR, 2017).

  • Porebski, B. T. et al. Rapid discovery of high-affinity antibodies via massively parallel sequencing, ribosome display and affinity screening. Nat. Biomed. Eng. 8, 214–232 (2024).

    Google Scholar 

  • Otwinowski, J., McCandlish, D. M. & Plotkin, J. B. Inferring the shape of global epistasis. Proc. Natl Acad. Sci. USA 115, E7550–E7558 (2018).

    Google Scholar 

  • Case, M., Smith, M., Vinh, J. & Thurber, G. Machine learning to predict continuous protein properties from binary cell sorting data and map unseen sequence space. Proc. Natl Acad. Sci. USA 121, e2311726121 (2024).

    Google Scholar 

  • Teney, D., Lin, Y., Oh, S.J. & Abbasnejad, E. ID and OOD performance are sometimes inversely correlated on real‑worlddatasets. Adv. Neural Inf. Process. Syst. 36, 36661–36673 (2023).

    Google Scholar 

  • Miller, J.P. et al. Accuracy on the line: on the strong correlation between out-of-distribution and in-distribution generalization. In Proc. 38th International Conference on Machine Learning (eds Meila, M. & Zhang, T.) 7721–7735 (PMLR, 2021).

  • Adams, R. M., Kinney, J. B., Walczak, A. M. & Mora, T. Epistasis in a fitness landscape defined by antibody-antigen binding free energy. Cell Syst. 8, 86–93.e3 (2019).

    Google Scholar 

  • Starr, T. N. & Thornton, J. W. Epistasis in protein evolution. Protein Sci. 25, 1204–1218 (2016).

    Google Scholar 

  • Cocco, S., Posani, L. & Monasson, R. Minimal epistatic networks from integrated sequence and mutational protein data. Preprint at bioRxiv https://doi.org/10.1101/2023.09.25.559251 (2023).

  • Papadopoulou, I., Nguyen, A.-P., Weber, A. & Martínez, M. R. DECODE: a computational pipeline to discover T cell receptor binding rules. Bioinformatics 38, i246–i254 (2022).

    Google Scholar 

  • Shanehsazzadeh, A. et al. Unlocking de novo antibody design with generative artificial intelligence. Preprint at bioRxiv https://doi.org/10.1101/2023.01.08.523187 (2023).

  • Leben, D. Explainable AI as evidence of fair decisions. Front. Psychol. 14, 1069426 (2023).

    Google Scholar 

  • Yang, D., Singh, A., Wu, H. & Kroe-Barrett, R. Comparison of biosensor platforms in the evaluation of high affinity antibody-antigen binding kinetics. Anal. Biochem. 508, 78–96 (2016).

    Google Scholar 

  • Mason, D. M. et al. Optimization of therapeutic antibodies by predicting antigen specificity from antibody sequence via deep learning. Nat. Biomed. Eng. https://doi.org/10.1038/s41551-021-00699-9 (2021).

    Google Scholar 

  • Liu, G. et al. Antibody complementarity determining region design using high-capacity machine learning. Bioinformatics 36, 2126–2133 (2020).

    Google Scholar 

  • Pei, Q. et al. Breaking the barriers of data scarcity in drug-target affinity prediction. Brief. Bioinform. 24, bbad386 (2023).

    Google Scholar 

  • Hu, F., Jiang, J., Wang, D., Zhu, M. & Yin, P. Multi-PLI: interpretable multi-task deep learning model for unifying protein-ligand interaction datasets. J. Cheminform. 13, 30 (2021).

    Google Scholar 

  • Kulikova, A. V. et al. Two sequence- and two structure-based ML models have learned different aspects of protein biochemistry. Sci. Rep. 13, 13280 (2023).

    Google Scholar 

  • Freschlin, C. R., Fahlberg, S. A., Heinzelman, P. & Romero, P. A. Neural network extrapolation to distant regions of the protein fitness landscape. Nat. Commun. 15, 6405 (2024).

    Google Scholar 

  • Yang, Z. et al. Does negative sampling matter? A review with insights into its theory and applications. IEEE Trans. Pattern Anal. Mach. Intell. 46, 5692–5711 (2024).

    Google Scholar 

  • Biswas, S., Khimulya, G., Alley, E. C., Esvelt, K. M. & Church, G. M. Low-N protein engineering with data-efficient deep learning. Nat. Methods 18, 389–396 (2021).

    Google Scholar 

  • Fowler, D. M. & Fields, S. Deep mutational scanning: a new style of protein science. Nat. Methods 11, 801–807 (2014).

    Google Scholar 

  • Greiff, V. et al. Systems analysis reveals high genetic and antigen-driven predetermination of antibody repertoires throughout B cell development. Cell Rep. 19, 1467–1478 (2017).

    Google Scholar 

  • Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D. & Wilson, A.G. Averaging weights leads to wider optima and better generalization. Preprint at https://arxiv.org/abs/1803.05407 (2018).

  • Athiwaratkun, B., Finzi, M., Izmailov, P. & Wilson, A.G. There are many consistent explanations of unlabeled data: why you should average. Preprint at https://arxiv.org/abs/1806.05594 (2018).

  • Paszke, A. et al. PyTorch: an imperative style, high-performance deep learning library Preprint at https://arxiv.org/abs/1912.01703 (2019).

  • Lin, Z. et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123–1130 (2023).

    MathSciNet 

    Google Scholar 

  • Leem, J., Mitchell, L. S., Farmery, J. H. R., Barton, J. & Galson, J. D. Deciphering the language of antibodies using self-supervised learning. Patterns 3, 100513 (2022).

    Google Scholar 

  • Barton, J., Galson, J.D. & Leem, J. Enhancing antibody language models with structural information. Preprint at bioRxiv https://doi.org/10.1101/2023.12.12.569610 (2024).

  • Cock, P. J. A. et al. Biopython: freely available Python tools for computational molecular biology and bioinformatics. Bioinformatics 25, 1422–1423 (2009).

    Google Scholar 

  • Kokhlikyan, N. et al. Captum: a unified and generic model interpretability library for PyTorch. Preprint at https://arxiv.org/abs/2009.07896 (2020).

  • Sundararajan, M., Taly, A. & Yan, Q. Axiomatic attribution for deep networks. In Proc. 34th International Conference on Machine Learning (eds Precup, D. & Teh, Y. W.) 3319–3328 (PMLR, 2017).

  • Novakovsky, G., Dexter, N., Libbrecht, M. W., Wasserman, W. W. & Mostafavi, S. Obtaining genetics insights from deep learning via explainable artificial intelligence. Nat. Rev. Genet. 24, 125–137 (2023).

    Google Scholar 

  • Huang, D. et al. Weakly supervised learning of RNA modifications from low-resolution epitranscriptome data. Bioinformatics 37, i222–i230 (2021).

    Google Scholar 

  • Ancona, M., Ceolini, E., Öztireli, C. & Gross, M. Towards better understanding of gradient-based attribution methods for deep neural networks. Preprint at https://arxiv.org/abs/1711.06104 (2017).

  • Ursu, E. & Minnegalieva, A. Training data composition determines machine learning generalization and biological rule discovery. Zenodo https://doi.org/10.5281/zenodo.11191740 (2024).

    Google Scholar 

  • Porebski, B. Rapid discovery of high-affinity antibodies via massively parallel sequencing, ribosome display and affinity screening. Zenodo https://doi.org/10.5281/zenodo.8241732 (2023).



  • Source link

    Leave a Reply

    Your email address will not be published. Required fields are marked *