A family of large language models for materials research with insights into model adaptability in continued pretraining

Machine Learning


  • Farley, I. Crossref documentation. Crossref https://www.crossref.org/documentation/retrieve-metadata/rest-api (2020).

  • Krishnan, N.M.A., Kodamana, H. & Bhattoo, R. Machine Learning for Materials Discovery: Numerical Recipes and Practical Applications (Springer, 2024).

  • Venugopal, V. & Olivetti, E. MatKG: an autonomously generated knowledge graph in material science. Sci. Data 11, 217 (2024).

    Article 

    Google Scholar 

  • Miret, S. & Krishnan, N. M. A. Enabling large language models for real-world materials discovery. Nat. Mach. Intell. 7, 991–998 (2025).

    Article 

    Google Scholar 

  • Bubeck, S. et al. Sparks of artificial general intelligence: early experiments with GPT-4. Preprint at https://arxiv.org/abs/2303.12712 (2023).

  • Gupta, T., Zaki, M., Krishnan, N. M. A., & Krishnan, M. Mausam Matscibert: a materials domain language model for text mining and information extraction. npj Comput. Mater. 8, 102 (2022).

    Article 

    Google Scholar 

  • Schilling-Wilhelmi, M. et al. From text to insight: large language models for materials science data extraction. Chem. Soc. Rev. 54, 1125–1150 (2025).

    Article 

    Google Scholar 

  • Mysore, S. et al. The materials science procedural text corpus: annotating materials synthesis procedures with shallow semantic structures. In Proc. 13th Linguistic Annotation Workshop (eds Friedrich, A. et al.) 56–64 (ACL, 2019).

  • Antunes, L. M., Butler, K. T. & Grau-Crespo, R. Crystal structure generation with autoregressive large language modeling. Nat. Commun. 15, 10570 (2024).

    Article 

    Google Scholar 

  • Gruver, N. et al. Fine-tuned language models generate stable inorganic materials as text. GitHub https://github.com/facebookresearch/crystal-text-llm (2024).

  • Ding, Q., Miret, S. & Liu, B. Matexpert: Decomposing materials discovery by mimicking human experts. In Proc. AI for Accelerated Materials Design (NeurIPS, 2024).

  • Boiko, D. A., MacKnight, R., Kline, B. & Gomes, G. Autonomous chemical research with large language models. Nature 624, 570–578 (2023).

    Article 

    Google Scholar 

  • Bran, A. M. et al. Augmenting large language models with chemistry tools. Nat. Mach. Intell. 6, 525–535 (2024).

    Article 

    Google Scholar 

  • Sim, M. et al. Chemos 2.0: an orchestration architecture for chemical self-driving laboratories. Matter 7, 2959–2977 (2024).

    Article 

    Google Scholar 

  • Zaki, M. et al. Mascqa: investigating materials science knowledge of large language models. Digital Discov. 3, 313–327 (2024).

    Article 

    Google Scholar 

  • White, A. D. et al. Assessment of chemistry knowledge in large language models that generate code. Digital Discov. 2, 368–376 (2023).

    Article 

    Google Scholar 

  • Dagdelen, J. et al. Structured information extraction from scientific text with large language models. Nat. Commun. 15, 1418 (2024).

    Article 

    Google Scholar 

  • Sayeed, H. M., Smallwood, W., Baird, S. G. & Sparks, T. D. Nlp meets materials science: quantifying the presentation of materials data in literature. Matter 7, 723–727 (2024).

    Article 

    Google Scholar 

  • Alampara, N., Miret, S. & Jablonka, K. M. Mattext: do language models need more than text and scale for materials modeling? In Proc. AI for Accelerated Materials Design (OpenReview.net, 2024).

  • Zimmermann, Y. et al. 32 examples of LLM applications in materials science and chemistry: towards automation, assistants, agents, and accelerated scientific discovery. Mach. Learn. Sci. Technol. 6, 030701 (2025).

    Article 

    Google Scholar 

  • Mirza, A. et al. Are large language models superhuman chemists? Nat. Chem. 17, 984–985 (2025).

    Article 

    Google Scholar 

  • Zhang, H., Song, Y., Hou, Z., Miret, S. & Liu, B. Honeycomb: a flexible LLM-based agent system for materials science. In Proc. Findings of the Association for Computational Linguistics: EMNLP (eds Al-Onaizan, Y. et al.) 3369–3382 (2024).

  • Hira, K. et al. Reconstructing the materials tetrahedron: challenges in materials information extraction. Digit. Discov. 3, 1021–1037 (2024).

    Article 

    Google Scholar 

  • Alampara, N. et al. Probing the limitations of multimodal language models for chemistry and materials research. Nat. Comput. Sci. 5, 952–961 (2025).

    Article 

    Google Scholar 

  • Song, Y., Miret, S., Zhang, H., and Liu, B. Honeybee: progressive instruction finetuning of large language models for materials science. In Proc. Findings of the Association for Computational Linguistics EMNLP (eds Bouamor, H. et al.) 5724–5739 (ACL, 2023).

  • Sayeed, H. M., Mohanty, T. & Sparks, T. D. Annotating materials science text: a semi-automated approach for crafting outputs with Gemini Pro. Integr. Mater. Manuf. Innov. 13, 445–452 (2024).

    Article 

    Google Scholar 

  • Circi, D., Khalighinejad, G., Chen, A., Dhingra, B. & Brinson, L. C. How well do large language models understand tables in materials science? Integr. Mater. Manuf. Innov. 13, 669–687 (2024).

    Article 

    Google Scholar 

  • Baird, S. G., Sayeed, H. M., Montoya, J. & Sparks, T. D. matbench-genmetrics: a python library for benchmarking crystal structure generative models using time-based splits of materials project structures. J. Open Source Softw. 9, 5618 (2024).

    Article 

    Google Scholar 

  • Touvron, H. et al. Llama 2: open foundation and fine-tuned chat models. Preprint at https://arxiv.org/abs/2307.09288 (2023).

  • Grattafiori, A. et al. The llama 3 herd of models. Preprint at https://arxiv.org/abs/2407.21783 (2024).

  • Mukherjee, S. et al. Orca: progressive learning from complex explanation traces of GPT-4. Preprint at https://arxiv.org/abs/2306.02707 (2023).

  • Hendrycks, D. et al. Measuring mathematical problem solving with the math dataset. In Proc. Neural Information Processing Systems Track on Datasets and Benchmarks Vol. 1 (eds Vanschoren, J. & Yeung, S.) (NeurIPS, 2021).

  • Song, Y., Miret, S. & Liu, B. MatSci-NLP: evaluating scientific language models on materials science language tasks using text-to-schema modeling. In Proc. 61st Annual Meeting of the Association for Computational Linguistics Vol. 1 (eds Rogers, A. et al.) 3621–3639 (ACL, 2023).

  • Song, Y., Miret, S., Zhang, H. & Liu, B. HoneyBee: progressive instruction finetuning of large language models for materials science. In Proc. Findings of the Association for Computational Linguistics EMNLP (eds Bouamor, H. et al.) 5724–5739 (ACL, 2023).

  • Siriwardhana, S. et al. Domain adaptation of llama3-70b-instruct through continual pre-training and model merging: a comprehensive evaluation. Preprint at https://arxiv.org/abs/2406.14971v1 (2024).

  • Öncel, F. et al. Adaptation odyssey in LLMs: why does additional pretraining sometimes fail to improve? In Proc. Conference on Empirical Methods in Natural Language Processing (eds Al-Onaizan, Y. et al.) 19834–19843 (ACL, 2024).

  • Polak, M. P. & Morgan, D. Extracting accurate materials data from research papers with conversational language models and prompt engineering. Nat. Commun. 15, 1569 (2024).

    Article 

    Google Scholar 

  • Gupta, T. et al. DiSCoMaT: distantly supervised composition extraction from tables in materials science articles. In Proc. 61st Annual Meeting of the Association for Computational Linguistics Vol. 1 (eds Rogers, A. et al.) 13465–13483 (ACL, 2023).

  • Springer, J. M. et al. Overtrained language models are harder to fine-tune. Preprint at https://arxiv.org/abs/2503.19206 (2025).

  • Dhruv, A et al. LLaMAT downstream evaluation dashboard. GitHub https://github.com/M3RG-IITD/llamat/tree/main/visualizations (2026).

  • Dhruv, A. et al. LLaMAT agent suite. GitHub http://github.com/M3RG-IITD/llamat/tree/main/agent (2026).

  • Ahlawat, D. et al. M3rg-iitd/llamat. GitHub https://github.com/M3RG-IITD/llamat (2025).

  • Siron, M. et al. Lemat-bulk: aggregating, and de-duplicating quantum chemistry materials databases. In Proc. AI for Accelerated Materials Design (ICLR, 2025).

  • Batatia, I. et al. A foundation model for atomistic materials chemistry. J. Chem. Phys. 163, 184110 (2025).

    Article 

    Google Scholar 

  • Wood, B. M. et al. Uma: a family of universal models for atoms. Preprint at https://arxiv.org/abs/2506.23971 (2025).

  • Rhodes, B. et al. Orb-v3: atomistic simulation at scale. Preprint at https://arxiv.org/abs/2504.06231 (2025).

  • Ahlawat, D. et al. Llamat. Zenodo https://zenodo.org/records/17251959 (2025).

  • Li, H., Xu, Z., Taylor, G., Studer, C. & Goldstein, T. Visualizing the loss landscape of neural nets. In Proc. Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems (eds Bengio, S. et al.) 6391–6401 (NeurIPS, 2018).

  • Science, health and medical journals, full text articles and books. ScienceDirect https://www.sciencedirect.com (2026).

  • APIs for research papers. Springer Nature Developer Portal https://dev.springernature.com (2026).

  • togethercomputer/RedPajama-Data-1T datasets at Hugging Face. Hugging Face (2024).

  • Ganose, A. M. & Jain, A. Robocrystallographer: automated crystal structure text descriptions and analysis. MRS Commun. 9, 874–881 (2019).

    Article 

    Google Scholar 

  • Jain, A. et al. in Handbook of Materials Modeling: Methods: Theory and Modeling 1751–1784 (2020).

  • Merchant, A. et al. Scaling deep learning for materials discovery. Nature 624, 80–85 (2023).

    Article 

    Google Scholar 

  • Downs, R. T. & Hall-Wallace, M. The American mineralogist crystal structure database. Am. Mineral. 88, 247–250 (2003).

    Google Scholar 

  • Hutagalung, S. D. Materials Science and Technology (IntechOpen, 2012).

  • Song, Y., Miret, S. & Liu, B. MatSci-NLP: evaluating scientific language models on materials science language tasks using text-to-schema modeling. Preprint at (2023).

  • Chen, Z. et al. Meditron-70b: scaling medical pretraining for large language models. Preprint at https://arxiv.org/abs/2311.16079 (2023).

  • Friedrich, A. et al. The SOFC-Exp corpus and neural approaches to information extraction in the materials science domain. In Proc. 58th Annual Meeting of the Association for Computational Linguistics (eds Jurafsky, D. et al.) 1255–1268 (ACL, 2020).

  • Yamaguchi, K., Asahi, R., and Sasaki, Y. Sc-comics: a superconductivity corpus for materials informatics. In Proc. 12th Language Resources and Evaluation Conference pages (eds Calzolari, N. et al.) 6753–6760 (ELRA, 2020).

  • Venugopal, V. et al. Looking through glass: knowledge discovery from materials science literature using natural language processing. Patterns 2, 100290 (2021).

  • Wang, Z. et al. Ulsa: unified language of synthesis actions for the representation of inorganic synthesis protocols. Digit. Discov. 1, 313–324 (2022).

    Article 

    Google Scholar 

  • Weston, L. et al. Named entity recognition and normalization applied to large-scale information extraction from the materials science literature. J. Chem. Inf. Model. 59, 3692–3702 (2019).

    Article 

    Google Scholar 

  • Dey, N. et al. Cerebras-gpt: open compute-optimal language models trained on the cerebras wafer-scale cluster. Preprint at https://arxiv.org/abs/2304.03208 (2023).

  • Xie, T., Fu, X., Ganea, O.-E., Barzilay, R. & Jaakkola, T. S. Crystal diffusion variational autoencoder for periodic material generation. In Proc. International Conference on Learning Representations (NeurIPS, 2022).

  • Jiao, R. et al. Crystal structure prediction by joint equivariant diffusion. Adv. Neural Inf. Process. Syst. 36, (2024).

  • Levy, D. et al. Symmcd: symmetry-preserving crystal generation with diffusion models. In Proc. AI for Accelerated Materials Design (NeurIPS, 2024).

  • Miller, B. K., Chen, R. T. Q., Sriram, A. & Wood, B. M. FlowMM: generating materials with Riemannian flow matching. In Proc. 41st International Conference on Machine Learning 35664–35686 (JMLR, 2024).

  • Lee, K. L. K. et al. Matsciml: a broad, multi-task benchmark for solid-state materials modeling. Preprint at https://arxiv.org/abs/2309.05934 (2023).

  • Duval, A. et al. A hitchhiker’s guide to geometric gnns for 3D atomic systems. Preprint at https://arxiv.org/abs/2312.07511 (2023).

  • Miret, S. et al. The open matsci ML toolkit: a flexible framework for machine learning in materials science. Trans. Mach. Learn. Res. 3, 759–768 (2023).

    Google Scholar 

  • Bihani, V. et al. Egraffbench: evaluation of equivariant graph neural network force fields for atomistic simulations. Digit. Discov. 3, 759–768 (2024).

    Article 

    Google Scholar 

  • Davies, D. W. et al. Smact: semiconducting materials by analogy and chemical theory. J. Open Source Softw. 4, 1361 (2019).

    Article 

    Google Scholar 

  • Chen, C. & Ong, S. P. A universal graph deep learning interatomic potential for the periodic table. Nat. Comput. Sci. 2, 718–728 (2022).

    Article 

    Google Scholar 

  • Sengupta, A. et al. Robust and efficient fine-tuning of LLMs with Bayesian reparameterization of low-rank adaptation. Preprint at https://arxiv.org/abs/2411.04358 (2024).

  • Mandal, I. et al. Evaluating large language model agents for automation of atomic force microscopy. Nat. Commun. 16, 9104 (2025).

    Article 

    Google Scholar 

  • Miglani, V., Yang, A., Markosyan, A. H., Garcia-Olano, D. &Kokhlikyan, N. Using captum to explain generative language models. Preprint at https://arxiv.org/abs/2312.05491 (2023).

  • Dhruv, A. et al. M3RG-IITD/llamat. GitHub https://github.com/M3RG-IITD/llamat/blob/main/src/ft_eval.py (2026).



  • Source link