An unprecedented dataset of molecular simulations to train released AI models

Machine Learning


Los Alamos contributes to unprecedented datasets and trains AI models | News Sways

In this representation, created with architect software, lanthanum, an example of the open molecule 2025 dataset, is surrounded by a variety of binding molecules. Lanthanum alloys are used in battery and hydrogen gas applications. Credit: Los Alamos National Laboratory

Meta, a joint effort between Lawrence Berkeley National Laboratory and Los Alamos National Laboratory leverages Los Alamos' expertise in building tools for molecular screening capabilities. The release of Open Molecules 2025, an unprecedented dataset for molecular simulation, can accelerate machine learning opportunities to transform research in areas such as biology, materials science, and energy technology.

The dataset is displayed in arxiv Preprint server.

“The outrageous part of molecular design was the extreme computational costs required to achieve quantum chemistry level accuracy,” said Michael G. Taylor, a researcher at Los Alamos and project member.

“To train machine learning models that can be accurate at quantum chemistry levels, a huge amount of diverse and effective training data is required. Open Molecules 2025 Bridge 2025 This gap bridges the dataset of density density functional theory calculations that can be used to train machine learning models accurately for any type of chemical challenge.”

Datasets are key to unlocking the use of machine learning possibilities for chemical applications, such as designing new drugs to combat battery cells to store disease and energy.

The adoption of density-functional theory calculations in the dataset allows for an accurate, atom level understanding of molecular behavior and interactions. The unique software designed by Taylor played a key role in the ability of Open Molecules 2025 to achieve their goals.

New software helps you build datasets

To help execute calculations and build datasets, the collaboration leveraged the capabilities of the architectural software designed by Taylor. Architector is cutting-edge software for predicting the 3D structure of metal complexes.

Metal complexes are chemicals in which the central metal atom binds to other molecules or arrays of atoms, representing important chemistry related to biology and materials science applications.

The architects employed in the theoretical departments of Taylor and Labs are primarily applied to the “F-block” element. These are lanthanides such as cerium and ytterbium, and actinides such as thorium and uranium.

F-block elements often contain many elements known as rare earth elements. It is valuable for a variety of industrial purposes, including high-tech applications such as telecommunications, imaging, and data storage.

Metal complexes represent the important classes of chemical classes investigated in the Open Molecule 2025 dataset. Other classes include ionic molecules such as proteins and RNAs, small molecules that can be the basis for drug discovery, and electrolyte metals surrounded by different solvents. Taylor estimates that chemistry investigated by the architects represents up to a third of the entire dataset.

Investing in basic chemistry knowledge

META has appointed its enormous computing power to perform density functional theory calculations. Considering only the rare earth molecule simulations that can be achieved, the Open Molecule 2025 project brought about approximately 20,000 structural data for each of the 17 rare earth elements.

The next largest dataset available in the literature has a total of about 1,000 structures per rare earth element.

Using the invaluable data generated, other machine learning models can now be trained in minutes and cost. Datasets can lead to pre-trained basic models that can be fine-tuned with minimal additional data in the area of ​​interest.

The efforts of Open Molecules 2025, including early machine learning models trained with data, are publicly available and provide researchers with the ability to use data and models relevant to their research.

“Chemical design often summarises in predicting the properties of new chemicals with minimal information and computational costs,” Taylor said.

“With this dataset, the ability to train machine learning models to perform its predictive work is potentially transformative for scientific discovery.”

detail:
Daniel S. Levine et al, The Open Molecules 2025 (Omol25) dataset, evaluation, and models; arxiv (2025). doi:10.48550/arxiv.2505.08762

Journal Information:
arxiv

Provided by Los Alamos National Laboratory

Quote: An unprecedented dataset of molecular simulations for training AI models released on June 13, 2025 https://phys.org/news/2025-06-unpreceded-dataset-molecular-simulations-ai.html

This document is subject to copyright. Apart from fair transactions for private research or research purposes, there is no part that is reproduced without written permission. Content is provided with information only.





Source link

Leave a Reply

Your email address will not be published. Required fields are marked *