
- Dataset volume
- ~500 TB
- DFT calculations
- >4M
- First release
- 4M random split
1FAIR at Meta, 2Carnegie Mellon University, 3Los Alamos National Laboratory, 4New York University, 5University of Cambridge, 6Stanford University, 7Prescient Design, Genentech, 8University of California-Berkeley, 9Lawrence Berkeley National Laboratory, 10Princeton University, 11Argonne National Laboratory, 12University of Chicago, 13Globus
* Corresponding author
Abstract
Background. The development of accurate machine learning models for molecular property prediction and materials design requires extensive high-quality training data. While databases of molecular structures and basic properties exist, comprehensive electronic structure data—including electronic densities, wavefunctions, and molecular orbitals—remain scarce at scale.
Objective. We present the OMol25 Electronic Structures dataset, an unprecedented open dataset of quantum chemical calculations designed to enable the development of next-generation physics-informed machine learning models for molecular chemistry and materials science.
Method. The dataset comprises raw density functional theory (DFT) outputs, electronic densities, wavefunctions, and molecular orbital information from over 4 million high-accuracy quantum chemical calculations performed on diverse molecular systems ranging from small organic molecules to large biomolecular complexes.
Impact. This dataset will enable researchers to develop improved partial charges, partial spins, and advanced electronic features for machine learning models, potentially accelerating discoveries in drug design, catalyst development, and energy materials.
01Introduction
We envision a future where researchers can rapidly design molecules and peptides to treat diseases, discover catalysts to revolutionize synthesis and manufacturing, identify the next electrolyte to store and transport energy to protect the grid, and more. But these breakthrough discoveries require data.
Data to train next-generation AI models and interatomic potentials. Data to push the boundaries of what's computationally possible in molecular chemistry and lead the world in AI for science. Data that captures the full complexity of chemical systems, from small organic molecules to massive biomolecular complexes.
02Dataset description
With our partners at Meta and Argonne Leadership Computing Facility (ALCF), we announce the OMol25 Electronic Structures dataset that includes ~500 TB of open molecular data. These data include the raw DFT outputs, electronic densities, wavefunctions, and molecular orbital information for over 4M high-accuracy quantum chemical calculations. We see this as a transformative opportunity to develop higher quality partial charges, partial spins, and advanced electronic features to unlock the next generation of physics-informed ML models.
The Materials Data Facility is proud to make these data available via the Eagle cluster at ALCF through a high-performance Globus endpoint. Given the dataset's unprecedented scale, we're first releasing all output data for a 4M random OMol25 split, with the full multi-petabyte dataset following based on community engagement.
03Research applications
This release supports new work across computational chemistry and machine learning.
01Train advanced ML models
Develop next-generation interatomic potentials and physics-informed models with unprecedented electronic structure data.
02Molecular property prediction
Build databases of calculated properties, partial charges, and descriptors for molecular screening.
03Catalyst discovery
Accelerate the discovery of novel catalysts for synthesis, manufacturing, and energy applications.
04Drug design
Leverage quantum chemical calculations to design molecules and peptides for treating diseases.
04Community and future directions
For this first release, the data are quite raw, and as-created by the Meta team. There's a significant opportunity for the community to build tools that simplify access to these data, allow data query and browsing, create databases of calculated properties and descriptors, and much more. We intend to work on these topics with all of you.
We can't wait to see what you can do with these data!
05Access the data
Files are served from the Eagle cluster at ALCF for high-performance transfer. At roughly half a petabyte, Globus is the recommended way to move them.
Access to this dataset requires a free Globus account and joining a permission group (link below). Due to the dataset's size (500TB), we recommend using Globus Connect Personal or Server for high-speed transfer.
06Acknowledgments
This release brings together data, infrastructure, and research expertise from MDF, Meta, ALCF, Globus, and university collaborators. MDF also acknowledges NIST and James Warren for their support of open materials-data infrastructure.
Key contributors
- Muhammed Shuaibi (Meta AI)
- Kyle Michel (Meta AI)
- Zachary Ulissi (Meta AI)
- Ian Foster (Argonne/UChicago)
- Kyle Chard (UChicago)
- Ben Blaiszik (Argonne)
References
- OMol25 preprintarXiv:2505.08762
Citation
Full citation will be provided upon official publication.
All Spotlight releases