AFLOW
High-throughput computational materials database with structures, properties, prototypes, and APIs. Best used for: - materials informatics - property lookup Category: Materials and catalysis.
Find public molecular property datasets, drug-discovery benchmarks, reaction records, synthesis datasets, spectra data, safety resources, and machine-learning-ready chemistry data.
Chemistry datasets are the foundation for molecular design, synthesis planning, assay modeling, spectra interpretation, safety lookup, and reproducible AI chemistry. This hub emphasizes reusable datasets with clear external links, citations, contributor acknowledgement, and enough context for researchers to decide whether a resource fits a project.
121 approved dataset entries are available for search indexing and community discovery.
High-throughput computational materials database with structures, properties, prototypes, and APIs. Best used for: - materials informatics - property lookup Category: Materials and catalysis.
Large conformational DFT dataset for organic molecules, foundational for transferable neural network potentials. Best used for: - machine-learned potentials - conformation energy prediction Category: Quantum chemistry.
High-accuracy coupled-cluster subset connected to ANI-1x for calibrating neural potentials and delta-learning models. Best used for: - high-accuracy energy prediction - delta learning Category: Quantum chemistry.
ANI active-learning dataset with energies and forces for chemically diverse organic conformations. Best used for: - active learning - potential training - force prediction Category: Quantum chemistry.
Quantum chemistry challenge dataset for molecular property prediction across many small organic molecules. Best used for: - benchmarking - property prediction Category: Quantum chemistry.
Curated aqueous solubility database with structures and experimental solubility values. Best used for: - solubility prediction - data curation Category: Physical chemistry.
BACE-1 inhibitor benchmark with activity labels for medicinal chemistry model evaluation. Best used for: - activity prediction - lead optimization Category: Bioactivity and drug discovery.
Blood-brain barrier penetration benchmark for small molecules. Best used for: - BBB prediction - CNS drug design Category: ADMET and toxicity.
Protein-ligand binding database with curated affinities, ligands, and complex structures. Best used for: - scoring - protein-ligand analysis Category: Protein-ligand and docking.
Public database of measured protein-ligand binding affinities from literature, patents, and submissions. Best used for: - affinity prediction - target-ligand SAR Category: Bioactivity and drug discovery.
Semi-manually curated protein-ligand interaction database for ligand binding and function annotation. Best used for: - binding-site annotation - protein function inference Category: Protein-ligand and docking.
Public archive of biomolecular NMR data, chemical shifts, restraints, and related experimental metadata. Best used for: - biomolecular NMR - chemical shift analysis Category: Spectra and analytical.
High-throughput Buchwald-Hartwig amination dataset used for reaction yield prediction and Bayesian optimization. Best used for: - yield modeling - reaction optimization Category: Reactions and synthesis.
C-N coupling optimization data used in reaction Bayesian optimization examples. Best used for: - reaction optimization - experimental design Category: Reactions and synthesis.
Comparative Assessment of Scoring Functions benchmark based on PDBbind core sets. Best used for: - scoring function assessment - docking validation Category: Protein-ligand and docking.
Collection of open natural product structures aggregated from many sources. Best used for: - natural product search - library design Category: Molecular libraries.
Benchmark suite for testing neural network potentials across conformers, drug-like molecules, and chemical diversity splits. Best used for: - potential benchmarking - out-of-distribution tests Category: Quantum chemistry.
Cytochrome P450 2D6 inhibition benchmark for metabolic liability prediction. Best used for: - CYP inhibition prediction - ADMET screening Category: ADMET and toxicity.
Cytochrome P450 3A4 inhibition benchmark for drug metabolism and interaction risk modeling. Best used for: - CYP inhibition prediction - drug interaction screening Category: ADMET and toxicity.
Curated dictionary and ontology of chemical entities of biological interest. Best used for: - identifier mapping - ontology annotation Category: Molecular libraries.
Manually curated database of bioactive molecules, assay measurements, targets, mechanisms, and medicinal chemistry literature. Best used for: - bioactivity modeling - target analysis - SAR Category: Bioactivity and drug discovery.
Benchmark for predicting clinical trial toxicity and FDA approval-related toxicity signals. Best used for: - clinical toxicity prediction - safety triage Category: ADMET and toxicity.
Cross-docked protein-ligand complexes used for pose prediction and deep docking model training. Best used for: - pose prediction - 3D generative modeling Category: Protein-ligand and docking.
Open collection of crystal structures in CIF format for organic, inorganic, metal-organic, and mineral compounds. Best used for: - structure lookup - crystallography Category: Materials and catalysis.
FDA-ranked drug-induced liver injury dataset for hepatotoxicity risk modeling. Best used for: - DILI prediction - safety assessment Category: ADMET and toxicity.
Directory of Useful Decoys, Enhanced: benchmark targets with actives and property-matched decoys. Best used for: - docking benchmark - virtual screening validation Category: Protein-ligand and docking.
Classic aqueous solubility benchmark for small organic molecules. Best used for: - logS prediction - property modeling Category: Physical chemistry.
Large-scale chemogenomics database combining ChEMBL and PubChem activity data for target prediction. Best used for: - target prediction - bioactivity modeling Category: Bioactivity and drug discovery.
Experimental and calculated hydration free energies for small molecules. Best used for: - solvation modeling - free-energy benchmarks Category: Physical chemistry.
Enumerated small organic molecules up to 13 heavy atoms for chemical space exploration. Best used for: - chemical space analysis - generative model evaluation Category: Molecular libraries.
Enumerated database of 166 billion organic molecules up to 17 heavy atoms. Best used for: - chemical space exploration - molecular generation Category: Molecular libraries.
Large conformer ensemble for drug-like molecules used to train and evaluate 3D molecular generative models. Best used for: - 3D generation - conformer prediction Category: Quantum chemistry.
Curated molecular conformer ensemble for QM9-scale molecules with geometry and conformational information. Best used for: - conformer generation - geometry prediction Category: Quantum chemistry.
Community platform and data resource for tandem mass spectrometry, molecular networking, and metabolomics annotation. Best used for: - metabolomics - MS/MS networking Category: Spectra and analytical.
Anti-HIV activity screen used as a challenging molecular classification benchmark. Best used for: - activity prediction - screening model evaluation Category: Bioactivity and drug discovery.
Curated metabolite resource with structures, biology, concentrations, and spectra for human metabolites. Best used for: - metabolite lookup - MS and NMR annotation Category: Spectra and analytical.
Molecular dynamics data for isomers with the same composition, useful for testing transfer across constitutional isomers. Best used for: - transfer learning - energy prediction - force prediction Category: Quantum chemistry.
DFT and force-field-ready materials datasets with structures, electronic, elastic, phonon, and other properties. Best used for: - materials screening - benchmarking Category: Materials and catalysis.
Curated lipid structure, classification, pathway, and MS resources for lipidomics. Best used for: - lipid annotation - lipidomics Category: Spectra and analytical.
Literature-derived protein-ligand virtual screening benchmark designed to reduce analogue bias. Best used for: - virtual screening - docking evaluation Category: Protein-ligand and docking.
Open natural products database connecting structures to organisms, taxonomy, and literature. Best used for: - chemodiversity analysis - natural product lookup Category: Molecular libraries.
Chemical component dictionary browser for PDB ligands, identifiers, structures, and dictionary metadata. Best used for: - ligand lookup - PDB component mapping Category: Protein-ligand and docking.
Octanol-water distribution coefficient benchmark commonly used for lipophilicity prediction. Best used for: - logD prediction - ADMET modeling Category: Physical chemistry.
Ab initio molecular dynamics trajectories for small molecules, commonly used for force-field and equivariant model benchmarks. Best used for: - force prediction - molecular dynamics surrogate modeling Category: Quantum chemistry.
Maximum Unbiased Validation benchmark designed for virtual screening tasks with reduced analogue bias. Best used for: - virtual screening - classification benchmarking Category: Bioactivity and drug discovery.
Public repository for high-quality mass spectra and compound annotations. Best used for: - MS identification - spectral matching Category: Spectra and analytical.
Open high-resolution mass spectral library with curated compound metadata and spectra. Best used for: - MS annotation - spectral library search Category: Spectra and analytical.
Repository for mass spectrometry datasets spanning proteomics, metabolomics, and small-molecule studies. Best used for: - raw MS data reuse - spectral reanalysis Category: Spectra and analytical.
Computed materials database with crystal structures, phase stability, electronic properties, and related data. Best used for: - materials screening - solid-state chemistry Category: Materials and catalysis.
Repository for metabolomics studies, raw data, metadata, and metabolite assignments. Best used for: - metabolomics data reuse - study lookup Category: Spectra and analytical.
MassBank of North America-derived spectral archive for metabolite and small-molecule MS/MS data. Best used for: - MS/MS library search - metabolite identification Category: Spectra and analytical.
Standard molecular machine-learning benchmark suite covering quantum, physical chemistry, biophysics, and physiology datasets. Best used for: - model comparison - molecular property prediction Category: Bioactivity and drug discovery.
Public web reference for thermochemical, ion energetics, IR, mass spectra, and gas chromatography data. Best used for: - reference lookup - spectral and thermochemical checks Category: Spectra and analytical.
Open database for organic structures and NMR chemical shifts, useful for NMR prediction and validation. Best used for: - NMR prediction - structure elucidation Category: Spectra and analytical.
Repository and archive for computational materials science data across electronic-structure codes. Best used for: - calculation reuse - materials ML Category: Materials and catalysis.
Curated database of microbial natural products with structures, taxonomy, and literature references. Best used for: - natural product discovery - dereplication Category: Molecular libraries.
Open database of DFT-computed formation energies and structures for inorganic materials. Best used for: - phase stability - materials discovery Category: Materials and catalysis.
Large DFT dataset of adsorbates on catalyst surfaces for training models that predict energies and relaxations. Best used for: - catalyst screening - surface ML Category: Materials and catalysis.
Oxide electrocatalyst dataset extending Open Catalyst benchmarks to OER-relevant surfaces and adsorbates. Best used for: - electrocatalyst modeling - surface energy prediction Category: Materials and catalysis.
Open schema and repository for chemical reaction data, including conditions, inputs, outcomes, and provenance. Best used for: - reaction data curation - yield modeling - condition analysis Category: Reactions and synthesis.
Large quantum chemistry dataset supporting OrbNet Denali models for fast electronic-structure-quality predictions. Best used for: - quantum ML - semiempirical correction Category: Quantum chemistry.
PubChem BioAssay subset curated by MoleculeNet for multi-task bioactivity prediction. Best used for: - multi-task learning - bioactivity prediction Category: Bioactivity and drug discovery.
Large quantum property benchmark for predicting DFT HOMO-LUMO gaps from molecular graphs. Best used for: - large-scale graph ML - quantum property prediction Category: Quantum chemistry.
Re-refined and rebuilt PDB structures that often improve geometry for structural analysis and docking preparation. Best used for: - structure preparation - model validation Category: Protein-ligand and docking.
Curated protein-ligand complex structures with experimentally measured binding affinities. Best used for: - binding affinity prediction - scoring function evaluation Category: Protein-ligand and docking.
Protein-ligand interaction dataset designed for leakage-aware benchmarking of structure-based ML models. Best used for: - structure-based ML - benchmarking Category: Protein-ligand and docking.
Large curated bioactivity dataset designed for reproducible machine learning in drug discovery. Best used for: - bioactivity prediction - benchmarking Category: Bioactivity and drug discovery.
Global archive of experimentally determined biomolecular structures, including protein-ligand complexes. Best used for: - structure lookup - docking preparation Category: Protein-ligand and docking.
NCBI chemical information resource with substance, compound, and BioAssay records for high-throughput and literature-linked activity data. Best used for: - assay lookup - screening analysis Category: Bioactivity and drug discovery.
Large computed quantum chemistry database built from PubChem compounds, including optimized structures and electronic properties. Best used for: - quantum property lookup - electronic structure ML Category: Quantum chemistry.
Small-molecule atomization energy benchmark derived from GDB molecules and widely used for early molecular representation studies. Best used for: - atomization energy prediction - molecular representations Category: Quantum chemistry.
Quantum chemistry dataset with equilibrium and non-equilibrium structures for small organic molecules, including energies and forces. Best used for: - force learning - property prediction Category: Quantum chemistry.
Extension of QM7 with multiple electronic properties for small molecules, useful for multi-target quantum property prediction. Best used for: - multi-property prediction - DFT benchmarks Category: Quantum chemistry.
Excited-state quantum chemistry benchmark for small molecules with calculated electronic spectra-related properties. Best used for: - excited-state prediction - spectroscopy ML Category: Quantum chemistry.
A benchmark of small organic molecules with DFT-computed geometries, energies, dipoles, orbital energies, thermochemistry, and related molecular properties. Best used for: - molecular property prediction - quantum ML - graph neural network...
Quantum mechanical properties and geometries for drug-like molecules from ChEMBL, useful for medicinal chemistry ML. Best used for: - drug-like quantum properties - 3D molecular ML Category: Quantum chemistry.
Database of Raman, infrared, X-ray diffraction, and chemistry data for minerals and related compounds. Best used for: - Raman identification - mineral chemistry Category: Spectra and analytical.
Reaction atom-mapping benchmark resources used with RXNMapper and USPTO reaction corpora. Best used for: - atom mapping - reaction model evaluation Category: Reactions and synthesis.
MS/MS spectral database for phytochemicals and plant metabolomics. Best used for: - plant metabolite annotation - MS/MS search Category: Spectra and analytical.
Cleaned and revised MD17 trajectories intended to reduce inconsistencies in molecular force-learning benchmarks. Best used for: - force prediction - model validation Category: Quantum chemistry.
Spectral Database for Organic Compounds with NMR, IR, Raman, ESR, and mass spectra. Best used for: - spectral lookup - compound identification Category: Spectra and analytical.
Drug side-effect dataset mapping marketed drugs to adverse event terms. Best used for: - side-effect prediction - drug safety analysis Category: ADMET and toxicity.
Quantum chemistry dataset spanning drug-like molecules, ions, peptides, and solvent-like systems for force-field and ML potential development. Best used for: - force-field fitting - ML potential training Category: Quantum chemistry.
Expert-classified reaction benchmark often paired with USPTO-50K retrosynthesis experiments. Best used for: - reaction classification - template evaluation Category: Reactions and synthesis.
Measured solubility benchmark originally used for community prediction comparisons. Best used for: - solubility benchmarking - model validation Category: Physical chemistry.
Natural product and natural-product-like compound database with structures and computed properties. Best used for: - natural product screening - library generation Category: Molecular libraries.
High-throughput Suzuki-Miyaura coupling dataset for data-driven reaction optimization. Best used for: - reaction optimization - yield prediction Category: Reactions and synthesis.
A curated platform of AI-ready datasets and benchmarks for therapeutic discovery and development. Best used for: - benchmarking - ADMET - drug discovery ML Category: Bioactivity and drug discovery.
Toxicology benchmark with nuclear receptor and stress-response pathway assay labels. Best used for: - toxicity prediction - safety screening Category: ADMET and toxicity.
Large collection of in vitro high-throughput toxicity assay measurements for environmental and drug-like chemicals. Best used for: - toxicity modeling - assay imputation Category: ADMET and toxicity.
Reaction condition prediction benchmark derived from patent reactions with reagents, catalysts, solvents, and temperature labels. Best used for: - condition recommendation - synthesis planning Category: Reactions and synthesis.
Large USPTO-derived benchmark frequently used for forward reaction prediction models. Best used for: - product prediction - reaction model training Category: Reactions and synthesis.
Curated 50k reaction subset from patent reactions, widely used for retrosynthesis and product prediction benchmarks. Best used for: - retrosynthesis - reaction prediction Category: Reactions and synthesis.
Large extracted reaction dataset from US patents, commonly used for synthesis planning and reaction model training. Best used for: - forward reaction prediction - retrosynthesis - condition prediction Category: Reactions and synthesis.
USPTO reaction subset preserving stereochemical information for retrosynthetic benchmarking. Best used for: - stereochemical retrosynthesis - reaction prediction Category: Reactions and synthesis.
Patent reaction yield datasets used for yield prediction and uncertainty-aware reaction modeling. Best used for: - yield prediction - reaction optimization Category: Reactions and synthesis.
Full patent reaction extraction used as a major public corpus for reaction informatics. Best used for: - reaction model pretraining - template extraction Category: Reactions and synthesis.
Ultra-large purchasable compound database prepared for ligand discovery and virtual screening. Best used for: - library screening - docking Category: Molecular libraries.
TDC hERG dataset family for potassium-channel liability prediction and cardiotoxicity triage. Best used for: - hERG prediction - cardiotoxicity screening Category: ADMET and toxicity.
Annotated binding-site database extracted from protein-ligand complexes in the PDB. Best used for: - binding-site analysis - docking setup Category: Protein-ligand and docking.
C2DB is a comprehensive database of computed properties for 2D materials, including structural, electronic, magnetic, and optical characteristics. It allows users to search and filter materials based on various criteria and provides detail...
A curated list of databases, datasets, books, and handbooks containing materials properties suitable for machine learning applications.
A curated list of superconductor databases and related resources provided by the IEEE Council on Superconductivity (CSC). Superconductor Databases & Resources
The 3DSC database is the first extensive collection of superconductors, including their critical temperature (Tc) and three-dimensional crystal structures. This repository provides access to the data and associated code. >9000 structures....
This repository contains the SuperCon-MTG dataset, including CSV files and descriptive files, used for accelerating the search for superconductors using machine learning. It also includes associated Python code. 13,731 compounds with compo...
A curated list of the most useful datasets in materials science and chemistry for training machine learning and AI foundation models. This includes experimental, computational, and literature-mined datasets, prioritizing open-access resour...
Materials-Related Databases These databases are sorted by the number of claimed data entries and by the existence of a method of programmatic access. Cloud Services - Globus - Globus Data Publication - Globus Catalog - Figshare - Center fo...
A dataset containing 546 uncorrelated carbon trajectories across various densities (1.0-3.5 gcm-3) and temperatures, capturing diverse chemical environments.
MatPES is a database of calculated potential energy surfaces for chemical reactions.
A dataset containing reaction mechanisms designed for training the Reactron model, a tool for predicting reaction outcomes.
The Materials Data Facility (MDF) is a platform for researchers to publish, discover, and access high-quality materials science datasets, promoting open data and accelerating scientific discovery.
OpenQDC is an open-source hub providing machine learning-ready quantum datasets. Users can explore existing datasets and contribute their own through tutorials.
Materials Cloud Discover provides access to curated research data sets with tailored visualizations, including databases for 3D and 2D crystal structures, pseudopotential libraries, and DFT implementation verification datasets.
A collection of materials and chemistry datasets for machine learning applications. This repository is archived and no longer maintained, with a recommendation to refer to the 'Awesome Materials & Chemistry Datasets' for current resources....
A curated list of known efforts in collecting and/or curating chemical and materials data, emphasizing the need for standards in data collection and sharing.
LeMat-Bulk is a unified dataset that combines Materials Project, OQMD, and Alexandria, containing over 5.3 million PBE-calculated materials. It also includes the largest collections of PBESol and SCAN functional calculations. The dataset s...
The Materials Project provides a free, open-access database of computed materials properties, enabling accelerated materials discovery.
The OMAT24 dataset, associated with ArXiv:2410.12771, is a materials science dataset from AI at Meta. It is licensed under CC BY 4.0.
The OQMD is a database containing DFT-calculated thermodynamic and structural properties for over 1.4 million materials. It offers tools for searching, querying, and analyzing material compositions, visualizing crystal structures, and acce...
A comprehensive materials database featuring crystal structures, geometries, phonon data, and benchmarks. Includes datasets for 3D, 2D, and 1D materials, optimized with PBE and PBEsol functionals, and offers access to generative models and...
OBELiX is a curated dataset containing 599 synthesized solid electrolyte materials, including their crystal structures and experimentally measured ionic conductivities. This dataset is designed to aid research in lithium solid-state batter...