Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K. BERT: pre-training of deep bidirectional transformers for language understanding. In Proc. 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) (eds Burstein, J. et al.) 4171–4186 (Association for Computational Linguistics, 2019).
Achiam, J. et al. GPT-4 technical report. Preprint at https://arxiv.org/abs/2303.08774 (2023).
Avsec, Ž. et al. Advancing regulatory variant effect prediction with AlphaGenome. Nature 649, 1206–1218 (2026).
Google Scholar
Brixi, G. et al. Genome modelling and design across all domains of life with Evo 2. Nature 652, 1349–1361 (2026).
Luo, X., Kang, X. & Schönhuth, A. Predicting the prevalence of complex genetic diseases from individual genotype profiles using capsule networks. Nat. Mach. Intell. 5, 114–125 (2023).
Google Scholar
Benegas, G., Albors, C., Aw, A. J., Ye, C. & Song, Y. S. A DNA language model based on multispecies alignment predicts the effects of genome-wide variants. Nat. Biotechnol. 43, 1960–1965 (2025).
Google Scholar
Xu, A. et al. SNPBag: a foundation model for multitask genome-scale SNP analysis. Preprint at Research Square https://doi.org/10.21203/rs.3.rs-7593414/v1 (2025).
Li, H. et al. BMFM-DNA: a SNP-aware DNA foundation model to capture variant effects. Preprint at https://arxiv.org/abs/2507.05265 (2025).
Gao, Z., Liu, Q., Zeng, W., Jiang, R. & Wong, W. H. EpiGePT: a pretrained transformer-based language model for context-specific human epigenomics. Genome Biol. 25, 310 (2024).
Google Scholar
Zeng, W., Guo, H., Liu, Q. & Wong, W. H. Improving polygenic prediction from whole-genome sequencing data by leveraging predicted epigenomic features. Proc. Natl Acad. Sci. USA 122, e2419202122 (2025).
Google Scholar
Zhou, H., Shrikumar, A. & Kundaje, A. Towards a better understanding of reverse-complement equivariance for deep learning models in genomics. In Proc. 16th Machine Learning in Computational Biology Meeting (eds Knowles, D. A. et al.) 1–33 (PMLR, 2022).
Ji, Y., Zhou, Z., Liu, H. & Davuluri, R. V. DNABERT: pre-trained bidirectional encoder representations from transformers model for DNA-language in genome. Bioinformatics 37, 2112–2120 (2021).
Google Scholar
Zhou, Z. et al. DNABERT-2: efficient foundation model and benchmark for multi-species genomes. In Proc. International Conference on Learning Representations https://openreview.net/forum?id=oMLQB4EZE1 (OpenReview, 2024).
Mallet, V. & Vert, J.-P. Reverse-complement equivariant networks for DNA sequences. Adv. Neural Inf. Process. Syst. 34, 13511–13523 (2021).
Schiff, Y. et al. Caduceus: bi-directional equivariant long-range DNA sequence modeling. In Proc. 41st International Conference on Machine Learning (eds Salakhutdinov, R. et al.) Vol. 235, 43632–43657 (PMLR, 2024).
Dalla-Torre, H. et al. Nucleotide transformer: building and evaluating robust foundation models for human genomics. Nat. Methods 22, 287–297 (2025).
Google Scholar
Nguyen, E. et al. HyenaDNA: long-range genomic sequence modeling at single nucleotide resolution. Adv. Neural Inf. Process. Syst. 36, 43177–43201 (2023).
Sanabria, M., Hirsch, J., Joubert, P. M. & Poetsch, A. R. DNA language model GROVER learns sequence context in the human genome. Nat. Mach. Intell. 6, 911–923 (2024).
Google Scholar
Press, O., Smith, N. A. & Lewis, M. Train short, test long: attention with linear biases enables input length extrapolation. In Proc. International Conference on Learning Representations https://openreview.net/forum?id=R8sQPpGCv0 (OpenReview, 2022).
Dao, T. FlashAttention-2: faster attention with better parallelism and work partitioning. In Proc. International Conference on Learning Representations (Curran Associates, 2024).
Lee, N. K., Tang, Z., Toneyan, S. & Koo, P. K. EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations. Genome Biol. 24, 105 (2023).
Google Scholar
Ma, M. Reverse-complement consistency for DNA language models. Preprint at https://arxiv.org/abs/2509.18529 (2025).
Duan, Q. et al. JanusDNA: a powerful bi-directional hybrid DNA foundation model. Adv. Neural Inf. Process. Syst. 38, 68791–68818 (2026).
Shrikumar, A., Greenside, P. & Kundaje, A. Reverse-complement parameter sharing improves deep learning models for genomics. Preprint at bioRxiv https://doi.org/10.1101/103663 (2017).
Choi, C. H. et al. DNA dynamically directs its own transcription initiation. Nucleic Acids Res 32, 1584–1590 (2004).
Google Scholar
Rohs, R. et al. The role of DNA shape in protein–DNA recognition. Nature 461, 1248–1253 (2009).
Google Scholar
Kabir, A. et al. DNA breathing integration with deep learning foundational model advances genome-wide binding prediction of human transcription factors. Nucleic Acids Res. 52, e91 (2024).
Gordân, R. et al. Genomic regions flanking E-box binding sites influence DNA binding specificity of bHLH transcription factors through DNA shape. Cell Rep. 3, 1093–1104 (2013).
Google Scholar
Chen, Y. et al. Structure of p53 binding to the BAX response element reveals DNA unwinding and compression to accommodate base-pair insertion. Nucleic Acids Res. 41, 8368–8376 (2013).
Google Scholar
Zhou, T. et al. Quantitative modeling of transcription factor binding specificities using DNA shape. Proc. Natl Acad. Sci. USA 112, 4654–4659 (2015).
Google Scholar
Mitra, R. et al. Geometric deep learning of protein–DNA binding specificity. Nat. Methods 21, 1674–1683 (2024).
Google Scholar
Adam, S., Klingel, V., Radde, N. E., Bashtrykov, P. & Jeltsch, A. On the accuracy of the epigenetic copy machine: comprehensive specificity analysis of the DNMT1 DNA methyltransferase. Nucleic Acids Res. 51, 6622–6633 (2023).
Google Scholar
Bashtrykov, P. et al. Specificity of Dnmt1 for methylation of hemimethylated CpG sites resides in its catalytic domain. Chem. Biol. 19, 572–578 (2012).
Google Scholar
Senadeera, D. C. et al. Dual branch VideoMamba with gated class token fusion for violence detection. Preprint at https://arxiv.org/abs/2506.03162 (2025).
Yu, T., Cheng, L., Khalitov, R., Olsson, E. B. & Yang, Z. Self-distillation improves self-supervised learning for DNA sequence inference. Neural Netw. 183, 106978 (2025).
Google Scholar
Yang, H. et al. HAD: hybrid architecture distillation outperforms teacher in genomic sequence modeling. Preprint at https://arxiv.org/abs/2505.20836 (2025).
Hu, J. et al. Comba: improving nonlinear RNNs with closed-loop control. Preprint at https://arxiv.org/abs/2505.12554 (2025).
Grešová, K., Martinek, V., Čechák, D., Šimeček, P. & Alexiou, P. Genomic Benchmarks: a collection of datasets for genomic sequence classification. BMC Genom. Data 24, 25 (2023).
Zhou, J. & Troyanskaya, O. G. Predicting effects of noncoding variants with deep learning-based sequence model. Nat. Methods 12, 931–934 (2015).
Google Scholar
Li, J., Pu, Y., Tang, J., Zou, Q. & Guo, F. DeepATT: a hybrid category attention neural network for identifying functional effects of DNA sequences. Brief. Bioinform. 22, bbaa159 (2021).
Google Scholar
Li, Z. et al. A novel interpretable deep learning-based computational framework designed synthetic enhancers with broad cross-species activity. Nucleic Acids Res. 52, 13447–13468 (2024).
Google Scholar
Cheng, W. et al. DNALONGBENCH: a benchmark suite for long-range DNA prediction tasks. Nat. Commun. 16, 10108 (2025).
Google Scholar
Avsec, Ž. et al. Effective gene expression prediction from sequence by integrating long-range interactions. Nat. Methods 18, 1196–1203 (2021).
Google Scholar
de Almeida, B. P., Reiter, F., Pagani, M. & Stark, A. DeepSTARR predicts enhancer activity from DNA sequence and enables the de novo design of synthetic enhancers. Nat. Genet. 54, 613–624 (2022).
Google Scholar
Gosai, S. J. et al. Machine-guided design of cell-type-targeting cis-regulatory elements. Nature 634, 1211–1220 (2024).
Google Scholar
Fishman, V. et al. GENA-LM: a family of open-source foundational DNA language models for long sequences. Nucleic Acids Res. 53, gkae1310 (2025).
Google Scholar
Linder, J., Srivastava, D., Yuan, H., Agarwal, V. & Kelley, D. R. Predicting RNA-seq coverage from DNA sequence as a unifying model of gene regulation. Nat. Genet. 57, 949–961 (2025).
Google Scholar
Ahmed, F. S., Aly, S. & Liu, X. EPI-Trans: an effective transformer-based deep learning model for enhancer promoter interaction prediction. BMC Bioinform. 25, 216 (2024).
Google Scholar
Feng, H. et al. Benchmarking DNA foundation models for genomic and genetic tasks. Nat. Commun. 16, 10780 (2025).
Google Scholar
Van Der Harst, P. & Verweij, N. Identification of 64 novel genetic loci provides an expanded view on the genetic architecture of coronary artery disease. Circ. Res. 122, 433–443 (2018).
Google Scholar
Medina, I. et al. Hck/Fgr kinase deficiency reduces plaque growth and stability by blunting monocyte recruitment and intraplaque motility. Circulation 132, 490–501 (2015).
Google Scholar
Bobryshev, Y. V. Monocyte recruitment and foam cell formation in atherosclerosis. Micron 37, 208–222 (2006).
Google Scholar
Li, J. et al. Novel insights: dynamic foam cells derived from the macrophage in atherosclerosis. J. Cell. Physiol. 236, 6154–6167 (2021).
Google Scholar
Heinz, S. et al. Simple combinations of lineage-determining transcription factors prime cis-regulatory elements required for macrophage and B cell identities. Mol. Cell 38, 576–589 (2010).
Google Scholar
Beltagy, I., Peters, M. E. & Cohan, A. Longformer: the long-document transformer. Preprint at https://arxiv.org/abs/2004.05150 (2020).
Schneider, V. A. et al. Evaluation of GRCh38 and de novo haploid genome assemblies demonstrates the enduring quality of the reference assembly. Genome Res. 27, 849–864 (2017).
Google Scholar
Nguyen, E. et al. Sequence modeling and design from molecular to genome scale with Evo. Science 386, eado9336 (2024).
Google Scholar
Su, J. et al. Roformer: enhanced transformer with rotary position embedding. Neurocomputing 568, 127063 (2024).
Google Scholar
Mou, L. et al. Natural language inference by tree-based convolution and heuristic matching. In Proc. 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) (eds Erk, K. et al.) 130–136 (Association for Computational Linguistics, 2016).
Upadhyaya, A., Nejdl, W. & Fisichella, M. Harnessing empathy and ethics for relevance detection and information categorization in climate and COVID-19 tweets. In Proc. 33rd ACM International Conference on Information and Knowledge Management 4091–4095 (ACM, 2024).
Zhang, K. et al. EATN: an efficient adaptive transfer network for aspect-level sentiment analysis. IEEE Trans. Knowl. Data Eng. 35, 377–389 (2021).
Morales-Brotons, D., Vogels, T. & Hendrikx, H. Exponential moving average of weights in deep learning: dynamics and benefits. Trans. Mach. Learn. Res. 2024, 1–27 (2024).
Yang, C. et al. CrossDNA: V1.1.0. Zenodo https://doi.org/10.5281/zenodo.19876436 (2026).
