Polymer research advances with OPoly26, a 6.57 million data point benchmark dataset

Machine Learning


Polymers form the building blocks of life and underpin countless technologies, yet their complex structures remain largely unknown to modern machine learning techniques. Meta's Daniel S. Levine of FAIR, Nicholas Liesen, and Lauren Chua of Lawrence Livermore National Laboratory, along with James Diffenderfer, Helgi Ingolfsson, and Matthew P. Kroonblawd, are addressing this gap by introducing the Open Polymers 2026 dataset, a substantial collection of more than 6.57 million calculations detailing the behavior of polymer systems. This achievement overcomes the significant computational challenges associated with modeling these large molecules and provides a resource that includes over 1.2 billion atoms and a wide range of polymer properties, including composition, structure, and environment. The research team demonstrated that incorporating this new data into machine learning training significantly increases the accuracy of polymer property predictions, paving the way for more versatile materials design and accelerating the development of universally applicable atomic models.

This dataset focuses on predicting radius of gyration and Flory index, which are important measures of polymer size and scaling, respectively. This approach focuses on generating a large and diverse set of polymer configurations using coarse-grained molecular dynamics simulations.

These simulations model the polymer chains as a series of beads, reducing computational costs while preserving important physical properties. A total of 10,000 unique polymer chains were simulated, each with 100 beads, representing a wide range of chemical compositions and chain stiffnesses. This extensive simulation data forms the basis for benchmarking machine learning models and includes a publicly available dataset of 10,000 polymer chains with corresponding radii of gyration and Flory index values. The dataset comes with an evaluation protocol and a baseline machine learning model, facilitating direct comparisons of different approaches. The researchers demonstrate the performance of several machine learning architectures on this dataset, establishing a new state-of-the-art technique for polymer property prediction. The OPoly26 dataset and associated tools are intended to facilitate future research in computational polymer science and enable the development of more accurate and efficient predictive models.

Molecular systems made up of repeating chemical units are the basis of life and drive advances in medicine, consumer products, and energy technology. Machine learning models have been trained on millions of quantum chemical simulations of materials and small molecules, but polymers have been largely excluded from datasets to date due to the computational cost of calculating accurate electronic structures. The central idea is to use a multistep process involving molecular dynamics (MD) simulations and reaction force field (AFIR) methods to create diverse datasets that go beyond static snapshots. MD simulations generate initial structures and allow exploration of structural space and dynamic behavior. A key innovation is the use of AFIR to capture chemical reactions and simulate bond breaking and formation, which are critical to generating diverse structures.

These generated structures are refined using density functional theory (DFT) or density functional strong binding (DTFB) calculations to obtain more accurate energies and shapes. Relevant substructures are extracted and capped with hydrogen atoms to prepare the final dataset. The workflow is designed for both polymers and lipids and employs a nonuniform sampling strategy during MD, with more frames acquired during the annealing step and selected based on dissimilarity to maximize structural diversity. AFIR is used in single-ended mode, specifying only the structure of the reactants and the bonds to be broken, allowing efficient exploration of reaction pathways.

A Universal Model for Atoms (UMA) force field trained on the OMol25 dataset calculates energy and forces during AFIR simulations. Post-processing involves creating protective zones around reactive bonds to maintain the local chemical environment and trimming the structure to maximum size to reduce computational cost. Significant efforts have been made to introduce charge diversity into the dataset, with approximately one-third of the structures being hydrogenated and one-third dehydrogenated. Various strategies are used to extract substructures from polymers and lipids to ensure chemically valid structures. The researchers assembled more than 6.57 million calculations representing more than 1.2 billion atoms to understand the diverse chemical properties of polymer systems. This extensive data set enables more accurate and transferable predictions of polymer properties, a challenge previously hampered by computational demands.

The research team demonstrated that incorporating the OPoly26 dataset into machine learning model training significantly improved energy predictions for polymers without negatively impacting the performance of predictions for smaller molecules. Importantly, this dataset complements existing molecular datasets and enables broad applicability in advanced models. Current research focuses on understanding how performance improvements are distributed across different polymer types and developing evaluations to assess the model's ability to predict local polymer structure and interactions. Future research will explore predicting properties of bulk polymers from experimental data, ultimately facilitating detailed computational studies of polymer behavior in applications such as fuel cells and materials upcycling. The OPoly26 dataset is publicly available to facilitate the collaborative development of more generalizable and accurate models of polymeric materials.



Source link