Hummer, A. M., Schneider, C., Chinery, L. & Deane, C. M. Investigating the volume and diversity of data needed for generalizable antibody–antigen ΔΔG prediction. Nat. Comput. Sci. https://doi.org/10.1038/s43588-025-00823-8 (2025).
Yang, R., Mao, J. & Chaudhari, P. Does the data induce capacity control in deep learning? In Proc. 39th International Conference on Machine Learning (ICML) (eds Chaudhuri, K. et al.) 25166–25193 (PMLR, 2022).
Alvarez-Melis, D. & Fusi, N. Geometric dataset distances via optimal transport. Adv. Neural Inf. Process. Syst. 33, 21428–21439 (2020).
Wang, T. & Isola, P. Understanding contrastive representation learning through alignment and uniformity on the hypersphere. In Proc. 37th International Conference on Machine Learning (ICML) (eds Daumé, H. & Singh, A.) 9929–9939 (PMLR, 2020).
Quiñonero-Candela, J., Sugiyama, M., Schwaighofer, A. & Lawrence, N. D. Dataset Shift in Machine Learning (The MIT Press, 2008).
Li, X. L., Liu, B. & Ng, S. K. Negative training data can be harmful to text classification. In Proc. 2010 Conference on Empirical Methods in Natural Language Processing (eds Li, H. & Màrquez, L.) 218–228 (Association for Computational Linguistics, 2010).
Wunsch, H., Kübler, S. & Cantrell, R. Instance sampling methods for pronoun resolution. In Proc. International Conference on Recent Advances in Natural Language Processing (RANLP-2009) (eds Angelova, G. & Mitkov, R.) 478–483 (Association for Computational Linguistics, 2009).
Saeidi, M., Kulkarni, R., Togia, T. & Sama, M. The effect of negative sampling strategy on capturing semantic similarity in document embeddings. In Proc. 2nd Workshop on Semantic Deep Learning (SemDeep-2) (eds Gromann, D. et al.) 1–8 (Association for Computational Linguistics, 2017).
Ben-Hur, A. & Noble, W. S. Choosing negative examples for the prediction of protein-protein interactions. BMC Bioinf. 7, S2 (2006).
Eid, F.-E., ElHefnawi, M. & Heath, L. S. DeNovo: virus-host sequence-based protein-protein interaction prediction. Bioinformatics 32, 1144–1150 (2016).
Tsukiyama, S., Hasan, M. M., Fujii, S. & Kurata, H. LSTM-PHV: prediction of human-virus protein-protein interactions by LSTM with word2vec. Brief. Bioinform. 22, bbab228 (2021).
Dens, C., Laukens, K., Bittremieux, W. & Meysman, P. The pitfalls of negative data bias for the T-cell epitope specificity challenge. Nat. Mach. Intell. https://doi.org/10.1038/s42256-023-00727-0 (2023).
Sidorczuk, K. et al. Benchmarks in antimicrobial peptide prediction are biased due to the selection of negative data. Brief. Bioinform. 23, bbac343 (2022).
Xu, L. et al. Negative sampling for contrastive representation learning: a review. Preprint at https://arxiv.org/abs/2206.00212 (2022).
Werner de Vargas, V., Schneider Aranda, J. A., Dos Santos Costa, R., da Silva Pereira, P. R. & Victória Barbosa, J. L. Imbalanced data preprocessing techniques for machine learning: a systematic mapping study. Knowl. Inf. Syst. 65, 31–57 (2023).
Loffredo, E., Pastore, M., Cocco, S. & Monasson, R. Restoring balance: principled under/oversampling of data for optimal classification. Preprint at https://arxiv.org/abs/2405.09535 (2024).
Akbar, R. et al. Progress and challenges for the machine learning-based design of fit-for-purpose monoclonal antibodies. MAbs 14, 2008790 (2022).
Greiff, V., Yaari, G. & Cowell, L. G. Mining adaptive immune receptor repertoires for biological and clinical information using machine learning. Curr. Opin. Syst. Biol. 24, 109–119 (2020).
Wilman, W. et al. Machine-designed biotherapeutics: opportunities, feasibility and advantages of deep learning in computational antibody discovery. Brief. Bioinform. 23, bbac267 (2022).
Khetan, R. et al. Current advances in biopharmaceutical informatics: guidelines, impact and challenges in the computational developability assessment of antibody therapeutics. MAbs 14, 2020082 (2022).
Fernández-Quintero, M. L. et al. Assessing developability early in the discovery process for novel biologics. MAbs 15, 2171248 (2023).
Chinery, L. et al. Baselining the buzz trastuzumab-HER2 affinity, and beyond. Preprint at bioRxiv https://doi.org/10.1101/2024.03.26.586756 (2024).
Ehling, R. A. et al. Synthetic coevolution reveals adaptive mutational trajectories of neutralizing antibodies and SARS-CoV-2. Preprint at bioRxiv https://doi.org/10.1101/2024.03.28.587189 (2024).
O’Donnell, T. J. et al. Reading the repertoire: Progress in adaptive immune receptor analysis using machine learning. Cell Syst 15, 1168–1189 (2024).
Schneider, C., Buchanan, A., Taddese, B. & Deane, C. M. DLAB: deep learning methods for structure-based virtual screening of antibodies. Bioinformatics 38, 377–383 (2022).
Krützfeldt, L.-M., Schubach, M. & Kircher, M. The impact of different negative training data on regulatory sequence predictions. PLoS ONE 15, e0237412 (2020).
Robert, P. A. et al. Unconstrained generation of synthetic antibody-antigen structures to guide machine learning methodology for antibody specificity prediction. Nat. Comput. Sci. 2, 845–865 (2022).
Montemurro, A., Jessen, L. E. & Nielsen, M. NetTCR-2.1: lessons and guidance on how to develop models for TCR specificity predictions. Front. Immunol. 13, 1055151 (2022).
Grazioli, F. et al. On TCR binding predictors failing to generalize to unseen peptides. Front. Immunol. 13, 1014256 (2022).
Deng, L. et al. Performance comparison of TCR-pMHC prediction tools reveals a strong data dependency. Front. Immunol. 14, 1128326 (2023).
Gao, Y., Gao, Y., Dong, K., Wu, S. & Liu, Q. Reply to: the pitfalls of negative data bias for the T-cell epitope specificity challenge. Nat. Mach. Intell. 5, 1063–1065 (2023).
Xu, J. L. & Davis, M. M. Diversity in the CDR3 region of V(H) is sufficient for most antibody specificities. Immunity 13, 37–45 (2000).
Davis, M. M. & Bjorkman, P. J. T-cell antigen receptor genes and T-cell recognition. Nature 334, 395–402 (1988).
Akbar, R. et al. A compact vocabulary of paratope-epitope interactions enables predictability of antibody-antigen binding. Cell Rep. 34, 108856 (2021).
Dunbar, J. et al. SAbDab: the structural antibody database. Nucleic Acids Res. 42, D1140–D1146 (2014).
Shugay, M. et al. VDJdb: a curated database of T-cell receptor sequences with known antigen specificity. Nucleic Acids Res. 46, D419–D427 (2018).
Tickotsky, N., Sagiv, T., Prilusky, J., Shifrut, E. & Friedman, N. McPAS-TCR: a manually curated catalogue of pathology-associated T cell receptor sequences. Bioinformatics 33, 2924–2929 (2017).
Swindells, M. B. et al. abYsis: integrated antibody sequence and structure-management, analysis, and prediction. J. Mol. Biol. 429, 356–364 (2017).
Raybould, M. I. J., Kovaltsuk, A., Marks, C. & Deane, C. M. CoV-AbDab: the coronavirus antibody database. Bioinformatics 37, 734–735 (2021).
Mahajan, S. et al. Epitope specific antibodies and T cell receptors in the immune epitope database. Front. Immunol. 9, 2688 (2018).
Chen, V. et al. Best practices for interpretable machine learning in computational biology. Preprint at bioRxiv https://doi.org/10.1101/2022.10.28.513978 (2022).
Sandve, G. K. & Greiff, V. Access to ground truth at unconstrained size makes simulated data as indispensable as experimental data for bioinformatics methods development and benchmarking. Bioinformatics 38, 4994–4996 (2022).
Khan, A. et al. Toward real-world automated antibody design with combinatorial Bayesian optimization.Cell Rep. Methods 3, 100374 (2023).
Zhang, C., Bzikadze, A. V., Safonova, Y. & Mirarab, S. A scalable model for simulating multi-round antibody evolution and benchmarking of clonal tree reconstruction methods. Front. Immunol. 13, 1014439 (2022).
Akbar, R. et al. In silico proof of principle of machine learning-based antibody design at unconstrained scale. MAbs 14, 2031482 (2022).
Sundararajan, M., Taly, A. & Yan, Q. Gradients of counterfactuals. Preprint at https://arxiv.org/abs/1611.02639 (2016).
Karim, M. R. et al. Explainable AI for bioinformatics: methods, tools and applications. Brief. Bioinform. 24, bbad236 (2023).
Shrikumar, A., Greenside, P. & Kundaje, A. Learning important features through propagating activation differences. In Proc. 34th International Conference on Machine Learning (eds Precup, D. & Teh, Y. W.) 3145–3153 (PMLR, 2017).
Porebski, B. T. et al. Rapid discovery of high-affinity antibodies via massively parallel sequencing, ribosome display and affinity screening. Nat. Biomed. Eng. 8, 214–232 (2024).
Otwinowski, J., McCandlish, D. M. & Plotkin, J. B. Inferring the shape of global epistasis. Proc. Natl Acad. Sci. USA 115, E7550–E7558 (2018).
Case, M., Smith, M., Vinh, J. & Thurber, G. Machine learning to predict continuous protein properties from binary cell sorting data and map unseen sequence space. Proc. Natl Acad. Sci. USA 121, e2311726121 (2024).
Teney, D., Lin, Y., Oh, S.J. & Abbasnejad, E. ID and OOD performance are sometimes inversely correlated on real‑worlddatasets. Adv. Neural Inf. Process. Syst. 36, 36661–36673 (2023).
Miller, J.P. et al. Accuracy on the line: on the strong correlation between out-of-distribution and in-distribution generalization. In Proc. 38th International Conference on Machine Learning (eds Meila, M. & Zhang, T.) 7721–7735 (PMLR, 2021).
Adams, R. M., Kinney, J. B., Walczak, A. M. & Mora, T. Epistasis in a fitness landscape defined by antibody-antigen binding free energy. Cell Syst. 8, 86–93.e3 (2019).
Starr, T. N. & Thornton, J. W. Epistasis in protein evolution. Protein Sci. 25, 1204–1218 (2016).
Cocco, S., Posani, L. & Monasson, R. Minimal epistatic networks from integrated sequence and mutational protein data. Preprint at bioRxiv https://doi.org/10.1101/2023.09.25.559251 (2023).
Papadopoulou, I., Nguyen, A.-P., Weber, A. & Martínez, M. R. DECODE: a computational pipeline to discover T cell receptor binding rules. Bioinformatics 38, i246–i254 (2022).
Shanehsazzadeh, A. et al. Unlocking de novo antibody design with generative artificial intelligence. Preprint at bioRxiv https://doi.org/10.1101/2023.01.08.523187 (2023).
Leben, D. Explainable AI as evidence of fair decisions. Front. Psychol. 14, 1069426 (2023).
Yang, D., Singh, A., Wu, H. & Kroe-Barrett, R. Comparison of biosensor platforms in the evaluation of high affinity antibody-antigen binding kinetics. Anal. Biochem. 508, 78–96 (2016).
Mason, D. M. et al. Optimization of therapeutic antibodies by predicting antigen specificity from antibody sequence via deep learning. Nat. Biomed. Eng. https://doi.org/10.1038/s41551-021-00699-9 (2021).
Liu, G. et al. Antibody complementarity determining region design using high-capacity machine learning. Bioinformatics 36, 2126–2133 (2020).
Pei, Q. et al. Breaking the barriers of data scarcity in drug-target affinity prediction. Brief. Bioinform. 24, bbad386 (2023).
Hu, F., Jiang, J., Wang, D., Zhu, M. & Yin, P. Multi-PLI: interpretable multi-task deep learning model for unifying protein-ligand interaction datasets. J. Cheminform. 13, 30 (2021).
Kulikova, A. V. et al. Two sequence- and two structure-based ML models have learned different aspects of protein biochemistry. Sci. Rep. 13, 13280 (2023).
Freschlin, C. R., Fahlberg, S. A., Heinzelman, P. & Romero, P. A. Neural network extrapolation to distant regions of the protein fitness landscape. Nat. Commun. 15, 6405 (2024).
Yang, Z. et al. Does negative sampling matter? A review with insights into its theory and applications. IEEE Trans. Pattern Anal. Mach. Intell. 46, 5692–5711 (2024).
Biswas, S., Khimulya, G., Alley, E. C., Esvelt, K. M. & Church, G. M. Low-N protein engineering with data-efficient deep learning. Nat. Methods 18, 389–396 (2021).
Fowler, D. M. & Fields, S. Deep mutational scanning: a new style of protein science. Nat. Methods 11, 801–807 (2014).
Greiff, V. et al. Systems analysis reveals high genetic and antigen-driven predetermination of antibody repertoires throughout B cell development. Cell Rep. 19, 1467–1478 (2017).
Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D. & Wilson, A.G. Averaging weights leads to wider optima and better generalization. Preprint at https://arxiv.org/abs/1803.05407 (2018).
Athiwaratkun, B., Finzi, M., Izmailov, P. & Wilson, A.G. There are many consistent explanations of unlabeled data: why you should average. Preprint at https://arxiv.org/abs/1806.05594 (2018).
Paszke, A. et al. PyTorch: an imperative style, high-performance deep learning library Preprint at https://arxiv.org/abs/1912.01703 (2019).
Lin, Z. et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123–1130 (2023).
Google Scholar
Leem, J., Mitchell, L. S., Farmery, J. H. R., Barton, J. & Galson, J. D. Deciphering the language of antibodies using self-supervised learning. Patterns 3, 100513 (2022).
Barton, J., Galson, J.D. & Leem, J. Enhancing antibody language models with structural information. Preprint at bioRxiv https://doi.org/10.1101/2023.12.12.569610 (2024).
Cock, P. J. A. et al. Biopython: freely available Python tools for computational molecular biology and bioinformatics. Bioinformatics 25, 1422–1423 (2009).
Kokhlikyan, N. et al. Captum: a unified and generic model interpretability library for PyTorch. Preprint at https://arxiv.org/abs/2009.07896 (2020).
Sundararajan, M., Taly, A. & Yan, Q. Axiomatic attribution for deep networks. In Proc. 34th International Conference on Machine Learning (eds Precup, D. & Teh, Y. W.) 3319–3328 (PMLR, 2017).
Novakovsky, G., Dexter, N., Libbrecht, M. W., Wasserman, W. W. & Mostafavi, S. Obtaining genetics insights from deep learning via explainable artificial intelligence. Nat. Rev. Genet. 24, 125–137 (2023).
Huang, D. et al. Weakly supervised learning of RNA modifications from low-resolution epitranscriptome data. Bioinformatics 37, i222–i230 (2021).
Ancona, M., Ceolini, E., Öztireli, C. & Gross, M. Towards better understanding of gradient-based attribution methods for deep neural networks. Preprint at https://arxiv.org/abs/1711.06104 (2017).
Ursu, E. & Minnegalieva, A. Training data composition determines machine learning generalization and biological rule discovery. Zenodo https://doi.org/10.5281/zenodo.11191740 (2024).
Porebski, B. Rapid discovery of high-affinity antibodies via massively parallel sequencing, ribosome display and affinity screening. Zenodo https://doi.org/10.5281/zenodo.8241732 (2023).
