All Spotlight releases

MDF SpotlightDataset article2025-10-07CC-BY-4.0

OMol25: A Large-Scale Electronic Structure Dataset for Accelerating Molecular Machine Learning

An open dataset of 500 TB comprising electronic densities, wavefunctions, and molecular orbital information from over 4 million high-accuracy density functional theory calculations

Scanning electron micrograph of a superconducting nanowire single-photon detector array
NASA/JPL · public domain
Illustrative specimen — SNSPD array, SEM~500 TB released
Dataset volume
~500 TB
DFT calculations
>4M
First release
4M random split
Daniel S. Levine1*,Muhammed Shuaibi1*,Evan Walter Clark Spotte-Smith2,Michael G. Taylor3,Muhammad R. Hasyim4,Kyle Michel1,Ilyes Batatia5,Gábor Csányi5,Misko Dzamba1,Peter Eastman6,Nathan C. Frey7,Xiang Fu1,Vahe Gharakhanyan1,Aditi S. Krishnapriyan8,9,Joshua A. Rackers7,Sanjeev Raja8,Ammar Rizvi1,Andrew S. Rosen10,Zachary Ulissi1,Santiago Vargas9,C. Lawrence Zitnick1*,Samuel M. Blau9*,Brandon M. Wood1*,Rachana Ananthakrishnan12,13,Rick Stevens11,Mike Papka11,Kyle Chard11,12,Ian Foster11,12,13,Ben Blaiszik11,12

1FAIR at Meta, 2Carnegie Mellon University, 3Los Alamos National Laboratory, 4New York University, 5University of Cambridge, 6Stanford University, 7Prescient Design, Genentech, 8University of California-Berkeley, 9Lawrence Berkeley National Laboratory, 10Princeton University, 11Argonne National Laboratory, 12University of Chicago, 13Globus

* Corresponding author

Abstract

Background. The development of accurate machine learning models for molecular property prediction and materials design requires extensive high-quality training data. While databases of molecular structures and basic properties exist, comprehensive electronic structure data—including electronic densities, wavefunctions, and molecular orbitals—remain scarce at scale.

Objective. We present the OMol25 Electronic Structures dataset, an unprecedented open dataset of quantum chemical calculations designed to enable the development of next-generation physics-informed machine learning models for molecular chemistry and materials science.

Method. The dataset comprises raw density functional theory (DFT) outputs, electronic densities, wavefunctions, and molecular orbital information from over 4 million high-accuracy quantum chemical calculations performed on diverse molecular systems ranging from small organic molecules to large biomolecular complexes.

Impact. This dataset will enable researchers to develop improved partial charges, partial spins, and advanced electronic features for machine learning models, potentially accelerating discoveries in drug design, catalyst development, and energy materials.

01Introduction

We envision a future where researchers can rapidly design molecules and peptides to treat diseases, discover catalysts to revolutionize synthesis and manufacturing, identify the next electrolyte to store and transport energy to protect the grid, and more. But these breakthrough discoveries require data.

Data to train next-generation AI models and interatomic potentials. Data to push the boundaries of what's computationally possible in molecular chemistry and lead the world in AI for science. Data that captures the full complexity of chemical systems, from small organic molecules to massive biomolecular complexes.

02Dataset description

With our partners at Meta and Argonne Leadership Computing Facility (ALCF), we announce the OMol25 Electronic Structures dataset that includes ~500 TB of open molecular data. These data include the raw DFT outputs, electronic densities, wavefunctions, and molecular orbital information for over 4M high-accuracy quantum chemical calculations. We see this as a transformative opportunity to develop higher quality partial charges, partial spins, and advanced electronic features to unlock the next generation of physics-informed ML models.

The Materials Data Facility is proud to make these data available via the Eagle cluster at ALCF through a high-performance Globus endpoint. Given the dataset's unprecedented scale, we're first releasing all output data for a 4M random OMol25 split, with the full multi-petabyte dataset following based on community engagement.

03Research applications

This release supports new work across computational chemistry and machine learning.

  1. 01Train advanced ML models

    Develop next-generation interatomic potentials and physics-informed models with unprecedented electronic structure data.

  2. 02Molecular property prediction

    Build databases of calculated properties, partial charges, and descriptors for molecular screening.

  3. 03Catalyst discovery

    Accelerate the discovery of novel catalysts for synthesis, manufacturing, and energy applications.

  4. 04Drug design

    Leverage quantum chemical calculations to design molecules and peptides for treating diseases.

04Community and future directions

For this first release, the data are quite raw, and as-created by the Meta team. There's a significant opportunity for the community to build tools that simplify access to these data, allow data query and browsing, create databases of calculated properties and descriptors, and much more. We intend to work on these topics with all of you.

We can't wait to see what you can do with these data!

05Access the data

Files are served from the Eagle cluster at ALCF for high-performance transfer. At roughly half a petabyte, Globus is the recommended way to move them.

Access to this dataset requires a free Globus account and joining a permission group (link below). Due to the dataset's size (500TB), we recommend using Globus Connect Personal or Server for high-speed transfer.

06Acknowledgments

This release brings together data, infrastructure, and research expertise from MDF, Meta, ALCF, Globus, and university collaborators. MDF also acknowledges NIST and James Warren for their support of open materials-data infrastructure.

Key contributors

  • Muhammed Shuaibi (Meta AI)
  • Kyle Michel (Meta AI)
  • Zachary Ulissi (Meta AI)
  • Ian Foster (Argonne/UChicago)
  • Kyle Chard (UChicago)
  • Ben Blaiszik (Argonne)

References

Citation

Full citation will be provided upon official publication.

All Spotlight releases
Materials Data Facility

The Materials Data Facility (MDF) empowers researchers to publish, discover, and access high-quality materials science datasets, accelerating scientific discovery through open data.

This work was performed under NIST financial assistance awards 70NANB14H012 and 70NANB19H005.

Supported by

NISTChiMaDUniversity of ChicagoArgonne National LaboratoryUniversity of Illinois

Cite MDF

Blaiszik, B., et al. "The Materials Data Facility: Data services to advance materials science research." JOM 68, no. 8 (2016): 2045-2052. doi:10.1007/s11837-016-2001-3

Blaiszik, B., et al. "A data ecosystem to support machine learning in materials science." MRS Communications 9, no. 4 (2019): 1125-1133. doi:10.1557/mrc.2019.118

© 2026 Materials Data Facility. All rights reserved.