Open Resources

Chemistry benchmarks

Benchmark pages help researchers compare models on shared tasks, reusable evaluation protocols, held-out datasets, and community challenge results.

Benchmarks for model comparison and reproducibility

Chemistry benchmarks make model claims easier to compare by tying tasks to datasets, metrics, splits, and evaluation protocols. This hub collects benchmark activities across molecular property prediction, generative molecule design, retrosynthesis, reaction prediction, spectra interpretation, docking, and virtual screening workflows.

Useful starting points

Related ChemistryAtlas pages

Approved resources

107 approved benchmark entries are available for search indexing and community discovery.

Benchmark · Docking and binding benchmarks

CASF-2016 ranking benchmark

CASF benchmark task for ranking protein-ligand complexes by affinity. Benchmark suite: CASF. Primary category: Docking and binding benchmarks. Leaderboard status: benchmark suite with published rankings.

Docking and binding benchmarksCASFchemistry benchmarkleaderboardPDBbind
Benchmark · Docking and binding benchmarks

CASF-2016 scoring benchmark

Comparative Assessment of Scoring Functions benchmark for protein-ligand scoring power. Benchmark suite: CASF. Primary category: Docking and binding benchmarks. Leaderboard status: benchmark suite with published rankings.

Docking and binding benchmarksCASFchemistry benchmarkleaderboardPDBbind
Benchmark · Docking and binding benchmarks

D3R Grand Challenge affinity ranking

Blind benchmark for ranking protein-ligand binding affinities across challenge targets. Benchmark suite: D3R Grand Challenges. Primary category: Docking and binding benchmarks. Leaderboard status: community blind challenge with public resu...

Docking and binding benchmarksD3R Grand Challengeschemistry benchmarkleaderboardD3R
Benchmark · Docking and binding benchmarks

D3R Grand Challenge free-energy prediction

Blind challenge benchmark for binding free-energy prediction. Benchmark suite: D3R Grand Challenges. Primary category: Docking and binding benchmarks. Leaderboard status: community blind challenge with public results.

Docking and binding benchmarksD3R Grand Challengeschemistry benchmarkleaderboardD3R
Benchmark · Docking and binding benchmarks

D3R Grand Challenge pose prediction

Blind docking benchmark focused on predicting protein-ligand binding poses. Benchmark suite: D3R Grand Challenges. Primary category: Docking and binding benchmarks. Leaderboard status: community blind challenge with public results.

Docking and binding benchmarksD3R Grand Challengeschemistry benchmarkleaderboardD3R
Benchmark · Virtual screening benchmarks

DUD-E virtual screening benchmark

Directory of Useful Decoys Enhanced benchmark for structure-based virtual screening. Benchmark suite: DUD-E. Primary category: Virtual screening benchmarks. Leaderboard status: benchmark dataset with extensive public comparisons.

Virtual screening benchmarksDUD-Echemistry benchmarkleaderboardvirtual screening
Benchmark · Generative molecular design benchmarks

GuacaMol

Benchmark suite for goal-directed and distribution-learning molecular generation. Benchmark suite: GuacaMol. Primary category: Generative molecular design benchmarks. Leaderboard status: published benchmark suite with reference leaderboard...

Generative molecular design benchmarksGuacaMolchemistry benchmarkleaderboardmolecular generation
Benchmark · Materials and spectra benchmarks

JARVIS-Leaderboard AI single-property tasks

NIST JARVIS leaderboard collection for AI models predicting material and molecular properties. Benchmark suite: JARVIS-Leaderboard. Primary category: Materials and spectra benchmarks. Leaderboard status: public JARVIS leaderboard.

Materials and spectra benchmarksJARVIS-Leaderboardchemistry benchmarkleaderboardJARVIS
Benchmark · Generative materials benchmarks

JARVIS-Leaderboard AtomGen tasks

JARVIS atomistic generation benchmark tasks for structure generation models. Benchmark suite: JARVIS-Leaderboard. Primary category: Generative materials benchmarks. Leaderboard status: public JARVIS leaderboard.

Generative materials benchmarksJARVIS-Leaderboardchemistry benchmarkleaderboardJARVIS
Benchmark · Catalysis and atomistic simulation

JARVIS-Leaderboard ML force-field tasks

JARVIS benchmark family for machine-learning force fields and atomistic simulations. Benchmark suite: JARVIS-Leaderboard. Primary category: Catalysis and atomistic simulation. Leaderboard status: public JARVIS leaderboard.

Catalysis and atomistic simulationJARVIS-Leaderboardchemistry benchmarkleaderboardJARVIS
Benchmark · Quantum chemistry benchmarks

JARVIS-Leaderboard electronic-structure tasks

JARVIS electronic-structure benchmarks for band gaps, energetics, and related properties. Benchmark suite: JARVIS-Leaderboard. Primary category: Quantum chemistry benchmarks. Leaderboard status: public JARVIS leaderboard.

Quantum chemistry benchmarksJARVIS-Leaderboardchemistry benchmarkleaderboardJARVIS
Benchmark · Spectra and analytical benchmarks

JARVIS-Leaderboard spectra tasks

JARVIS benchmarks for spectral-property prediction and analysis. Benchmark suite: JARVIS-Leaderboard. Primary category: Spectra and analytical benchmarks. Leaderboard status: public JARVIS leaderboard.

Spectra and analytical benchmarksJARVIS-Leaderboardchemistry benchmarkleaderboardJARVIS
Benchmark · Virtual screening benchmarks

LIT-PCBA virtual screening benchmark

Literature-derived PubChem BioAssay benchmark for virtual screening and target-specific ranking. Benchmark suite: LIT-PCBA. Primary category: Virtual screening benchmarks. Leaderboard status: benchmark dataset with public method comparison...

Virtual screening benchmarksLIT-PCBAchemistry benchmarkleaderboardvirtual screening
Benchmark · Generative molecular design benchmarks

MOSES

Molecular generation benchmark platform with datasets, metrics, and baseline generative models. Benchmark suite: MOSES. Primary category: Generative molecular design benchmarks. Leaderboard status: benchmark platform with baseline comparis...

Generative molecular design benchmarksMOSESchemistry benchmarkleaderboardmolecular generation
Benchmark · Materials and molecular property benchmarks

Matbench Discovery

Benchmark for materials discovery and stability prediction models. Benchmark suite: Matbench Discovery. Primary category: Materials and molecular property benchmarks. Leaderboard status: public benchmark suite and leaderboard.

Materials and molecular property benchmarksMatbench Discoverychemistry benchmarkleaderboardmaterials discovery
Benchmark · Materials and molecular property benchmarks

Matbench v0.1

Materials machine-learning benchmark suite, useful for ChemistryAtlas users working across molecular and materials discovery. Benchmark suite: Matbench. Primary category: Materials and molecular property benchmarks. Leaderboard status: pub...

Materials and molecular property benchmarksMatbenchchemistry benchmarkleaderboardmaterials
Benchmark · Bioactivity prediction benchmarks

MoleculeNet: BACE

MoleculeNet benchmark task for BACE inhibitor classification. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Bioactivity prediction bench...

Bioactivity prediction benchmarksMoleculeNetchemistry benchmarkleaderboardBACE
Benchmark · ADMET and toxicity leaderboards

MoleculeNet: BBBP

MoleculeNet benchmark task for blood-brain barrier penetration classification. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: ADMET and t...

ADMET and toxicity leaderboardsMoleculeNetchemistry benchmarkleaderboardBBBP
Benchmark · Materials and molecular property benchmarks

MoleculeNet: CEP

MoleculeNet benchmark task for organic photovoltaic candidate property prediction. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Materia...

Materials and molecular property benchmarksMoleculeNetchemistry benchmarkleaderboardCEP
Benchmark · ADMET and toxicity leaderboards

MoleculeNet: ClinTox

MoleculeNet benchmark task for clinical trial toxicity classification. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: ADMET and toxicity...

ADMET and toxicity leaderboardsMoleculeNetchemistry benchmarkleaderboardClinTox
Benchmark · Physical chemistry benchmarks

MoleculeNet: ESOL

MoleculeNet benchmark task for aqueous solubility regression. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Physical chemistry benchmark...

Physical chemistry benchmarksMoleculeNetchemistry benchmarkleaderboardESOL
Benchmark · Physical chemistry benchmarks

MoleculeNet: FreeSolv

MoleculeNet benchmark task for hydration free energy regression. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Physical chemistry benchm...

Physical chemistry benchmarksMoleculeNetchemistry benchmarkleaderboardFreeSolv
Benchmark · Bioactivity prediction benchmarks

MoleculeNet: HIV

MoleculeNet benchmark task for anti-HIV activity classification. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Bioactivity prediction be...

Bioactivity prediction benchmarksMoleculeNetchemistry benchmarkleaderboardHIV
Benchmark · Physical chemistry benchmarks

MoleculeNet: Lipophilicity

MoleculeNet benchmark task for octanol-water distribution regression. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Physical chemistry b...

Physical chemistry benchmarksMoleculeNetchemistry benchmarkleaderboardLipophilicity
Benchmark · Virtual screening benchmarks

MoleculeNet: MUV

MoleculeNet benchmark task for unbiased virtual-screening benchmark. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Virtual screening ben...

Virtual screening benchmarksMoleculeNetchemistry benchmarkleaderboardMUV
Benchmark · Bioactivity prediction benchmarks

MoleculeNet: PCBA

MoleculeNet benchmark task for PubChem BioAssay multi-task benchmark. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Bioactivity predicti...

Bioactivity prediction benchmarksMoleculeNetchemistry benchmarkleaderboardPCBA
Benchmark · Docking and binding benchmarks

MoleculeNet: PDBbind

MoleculeNet benchmark task for protein-ligand binding affinity prediction. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Docking and bin...

Docking and binding benchmarksMoleculeNetchemistry benchmarkleaderboardPDBbind
Benchmark · Quantum chemistry benchmarks

MoleculeNet: QM7

MoleculeNet benchmark task for quantum property regression. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Quantum chemistry benchmarks....

Quantum chemistry benchmarksMoleculeNetchemistry benchmarkleaderboardQM7
Benchmark · Quantum chemistry benchmarks

MoleculeNet: QM7b

MoleculeNet benchmark task for multi-property quantum chemistry regression. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Quantum chemis...

Quantum chemistry benchmarksMoleculeNetchemistry benchmarkleaderboardQM7b
Benchmark · Quantum chemistry benchmarks

MoleculeNet: QM8

MoleculeNet benchmark task for excited-state quantum-property prediction. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Quantum chemistr...

Quantum chemistry benchmarksMoleculeNetchemistry benchmarkleaderboardQM8
Benchmark · Quantum chemistry benchmarks

MoleculeNet: QM9

MoleculeNet benchmark task for small-molecule quantum-property prediction. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Quantum chemist...

Quantum chemistry benchmarksMoleculeNetchemistry benchmarkleaderboardQM9
Benchmark · ADMET and toxicity leaderboards

MoleculeNet: SIDER

MoleculeNet benchmark task for drug side-effect classification. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: ADMET and toxicity leaderb...

ADMET and toxicity leaderboardsMoleculeNetchemistry benchmarkleaderboardSIDER
Benchmark · ADMET and toxicity leaderboards

MoleculeNet: Tox21

MoleculeNet benchmark task for multi-task toxicity classification. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: ADMET and toxicity lead...

ADMET and toxicity leaderboardsMoleculeNetchemistry benchmarkleaderboardTox21
Benchmark · ADMET and toxicity leaderboards

MoleculeNet: ToxCast

MoleculeNet benchmark task for high-throughput toxicity assay classification. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: ADMET and to...

ADMET and toxicity leaderboardsMoleculeNetchemistry benchmarkleaderboardToxCast
Benchmark · Molecular graph and property prediction

OGB-LSC PCQM4Mv2

Large-scale molecular graph benchmark for HOMO-LUMO gap prediction from quantum chemistry labels. Benchmark suite: Open Graph Benchmark Large-Scale Challenge. Primary category: Molecular graph and property prediction. Leaderboard status: p...

Molecular graph and property predictionOpen Graph Benchmark Large-Scale Challengechemistry benchmarkleaderboardOGB
Benchmark · Catalysis and atomistic simulation

Open Catalyst OC20 IS2RE

Initial-structure-to-relaxed-energy task for catalyst adsorbate systems. Benchmark suite: Open Catalyst Project. Primary category: Catalysis and atomistic simulation. Leaderboard status: public Open Catalyst leaderboard.

Catalysis and atomistic simulationOpen Catalyst Projectchemistry benchmarkleaderboardOpen Catalyst
Benchmark · Catalysis and atomistic simulation

Open Catalyst OC20 IS2RS

Initial-structure-to-relaxed-structure benchmark for catalyst systems. Benchmark suite: Open Catalyst Project. Primary category: Catalysis and atomistic simulation. Leaderboard status: public Open Catalyst leaderboard.

Catalysis and atomistic simulationOpen Catalyst Projectchemistry benchmarkleaderboardOpen Catalyst
Benchmark · Catalysis and atomistic simulation

Open Catalyst OC20 S2EF

Structure-to-energy-and-forces benchmark for adsorbate-catalyst systems. Benchmark suite: Open Catalyst Project. Primary category: Catalysis and atomistic simulation. Leaderboard status: public Open Catalyst leaderboard.

Catalysis and atomistic simulationOpen Catalyst Projectchemistry benchmarkleaderboardOpen Catalyst
Benchmark · Catalysis and atomistic simulation

Open Catalyst OC22

Open Catalyst 2022 benchmark for oxide catalyst systems and atomistic ML. Benchmark suite: Open Catalyst Project. Primary category: Catalysis and atomistic simulation. Leaderboard status: Open Catalyst benchmark and challenge ecosystem.

Catalysis and atomistic simulationOpen Catalyst Projectchemistry benchmarkleaderboardOpen Catalyst
Benchmark · Catalysis and atomistic simulation

Open Catalyst OC25

Emerging Open Catalyst benchmark set for next-generation catalyst modeling. Benchmark suite: Open Catalyst Project. Primary category: Catalysis and atomistic simulation. Leaderboard status: emerging Open Catalyst benchmark.

Catalysis and atomistic simulationOpen Catalyst Projectchemistry benchmarkleaderboardOpen Catalyst
Benchmark · Spectra and analytical benchmarks

RamanBench

Benchmark and live leaderboard for Raman spectroscopy AI models. Benchmark suite: RamanBench. Primary category: Spectra and analytical benchmarks. Leaderboard status: published benchmark with live leaderboard.

Spectra and analytical benchmarksRamanBenchchemistry benchmarkleaderboardRaman spectroscopy
Benchmark · Docking and binding benchmarks

SAMPL host-guest binding challenge

Blind host-guest binding free-energy challenge for force-field, simulation, and ML methods. Benchmark suite: SAMPL Challenges. Primary category: Docking and binding benchmarks. Leaderboard status: community blind challenge with public resu...

Docking and binding benchmarksSAMPL Challengeschemistry benchmarkleaderboardSAMPL
Benchmark · Physical chemistry benchmarks

SAMPL hydration free-energy challenge

Blind challenge benchmark for hydration free-energy prediction using physical and ML methods. Benchmark suite: SAMPL Challenges. Primary category: Physical chemistry benchmarks. Leaderboard status: community blind challenge with public res...

Physical chemistry benchmarksSAMPL Challengeschemistry benchmarkleaderboardSAMPL
Benchmark · Physical chemistry benchmarks

SAMPL logP/logD challenge

Blind partition/distribution coefficient challenge for logP and logD prediction. Benchmark suite: SAMPL Challenges. Primary category: Physical chemistry benchmarks. Leaderboard status: community blind challenge with public results.

Physical chemistry benchmarksSAMPL Challengeschemistry benchmarkleaderboardSAMPL
Benchmark · Physical chemistry benchmarks

SAMPL pKa challenge

Blind pKa prediction challenge for small-molecule protonation and thermodynamics methods. Benchmark suite: SAMPL Challenges. Primary category: Physical chemistry benchmarks. Leaderboard status: community blind challenge with public results.

Physical chemistry benchmarksSAMPL Challengeschemistry benchmarkleaderboardSAMPL
Benchmark · ADMET and toxicity leaderboards

TDC ADMET: AMES

Official TDC ADMET Group leaderboard task for the AMES endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity leaderboard...

ADMET and toxicity leaderboardsTherapeutics Data Commons ADMET Groupchemistry benchmarkleaderboardTDC
Benchmark · ADMET and toxicity leaderboards

TDC ADMET: BBB_Martins

Official TDC ADMET Group leaderboard task for the BBB_Martins endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity lead...

ADMET and toxicity leaderboardsTherapeutics Data Commons ADMET Groupchemistry benchmarkleaderboardTDC
Benchmark · ADMET and toxicity leaderboards

TDC ADMET: Bioavailability_Ma

Official TDC ADMET Group leaderboard task for the Bioavailability_Ma endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxici...

ADMET and toxicity leaderboardsTherapeutics Data Commons ADMET Groupchemistry benchmarkleaderboardTDC
Benchmark · ADMET and toxicity leaderboards

TDC ADMET: CYP2C9_Substrate_CarbonMangels

Official TDC ADMET Group leaderboard task for the CYP2C9_Substrate_CarbonMangels endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADME...

ADMET and toxicity leaderboardsTherapeutics Data Commons ADMET Groupchemistry benchmarkleaderboardTDC
Benchmark · ADMET and toxicity leaderboards

TDC ADMET: CYP2C9_Veith

Official TDC ADMET Group leaderboard task for the CYP2C9_Veith endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity lea...

ADMET and toxicity leaderboardsTherapeutics Data Commons ADMET Groupchemistry benchmarkleaderboardTDC
Benchmark · ADMET and toxicity leaderboards

TDC ADMET: CYP2D6_Substrate_CarbonMangels

Official TDC ADMET Group leaderboard task for the CYP2D6_Substrate_CarbonMangels endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADME...

ADMET and toxicity leaderboardsTherapeutics Data Commons ADMET Groupchemistry benchmarkleaderboardTDC
Benchmark · ADMET and toxicity leaderboards

TDC ADMET: CYP2D6_Veith

Official TDC ADMET Group leaderboard task for the CYP2D6_Veith endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity lea...

ADMET and toxicity leaderboardsTherapeutics Data Commons ADMET Groupchemistry benchmarkleaderboardTDC
Benchmark · ADMET and toxicity leaderboards

TDC ADMET: CYP3A4_Substrate_CarbonMangels

Official TDC ADMET Group leaderboard task for the CYP3A4_Substrate_CarbonMangels endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADME...

ADMET and toxicity leaderboardsTherapeutics Data Commons ADMET Groupchemistry benchmarkleaderboardTDC
Benchmark · ADMET and toxicity leaderboards

TDC ADMET: CYP3A4_Veith

Official TDC ADMET Group leaderboard task for the CYP3A4_Veith endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity lea...

ADMET and toxicity leaderboardsTherapeutics Data Commons ADMET Groupchemistry benchmarkleaderboardTDC
Benchmark · ADMET and toxicity leaderboards

TDC ADMET: Caco2_Wang

Official TDC ADMET Group leaderboard task for the Caco2_Wang endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity leade...

ADMET and toxicity leaderboardsTherapeutics Data Commons ADMET Groupchemistry benchmarkleaderboardTDC
Benchmark · ADMET and toxicity leaderboards

TDC ADMET: Clearance_Hepatocyte_AZ

Official TDC ADMET Group leaderboard task for the Clearance_Hepatocyte_AZ endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and t...

ADMET and toxicity leaderboardsTherapeutics Data Commons ADMET Groupchemistry benchmarkleaderboardTDC
Benchmark · ADMET and toxicity leaderboards

TDC ADMET: Clearance_Microsome_AZ

Official TDC ADMET Group leaderboard task for the Clearance_Microsome_AZ endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and to...

ADMET and toxicity leaderboardsTherapeutics Data Commons ADMET Groupchemistry benchmarkleaderboardTDC
Benchmark · ADMET and toxicity leaderboards

TDC ADMET: DILI

Official TDC ADMET Group leaderboard task for the DILI endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity leaderboard...

ADMET and toxicity leaderboardsTherapeutics Data Commons ADMET Groupchemistry benchmarkleaderboardTDC
Benchmark · ADMET and toxicity leaderboards

TDC ADMET: HIA_Hou

Official TDC ADMET Group leaderboard task for the HIA_Hou endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity leaderbo...

ADMET and toxicity leaderboardsTherapeutics Data Commons ADMET Groupchemistry benchmarkleaderboardTDC
Benchmark · ADMET and toxicity leaderboards

TDC ADMET: Half_Life_Obach

Official TDC ADMET Group leaderboard task for the Half_Life_Obach endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity...

ADMET and toxicity leaderboardsTherapeutics Data Commons ADMET Groupchemistry benchmarkleaderboardTDC
Benchmark · ADMET and toxicity leaderboards

TDC ADMET: LD50_Zhu

Official TDC ADMET Group leaderboard task for the LD50_Zhu endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity leaderb...

ADMET and toxicity leaderboardsTherapeutics Data Commons ADMET Groupchemistry benchmarkleaderboardTDC
Benchmark · ADMET and toxicity leaderboards

TDC ADMET: Lipophilicity_AstraZeneca

Official TDC ADMET Group leaderboard task for the Lipophilicity_AstraZeneca endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and...

ADMET and toxicity leaderboardsTherapeutics Data Commons ADMET Groupchemistry benchmarkleaderboardTDC
Benchmark · ADMET and toxicity leaderboards

TDC ADMET: PPBR_AZ

Official TDC ADMET Group leaderboard task for the PPBR_AZ endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity leaderbo...

ADMET and toxicity leaderboardsTherapeutics Data Commons ADMET Groupchemistry benchmarkleaderboardTDC
Benchmark · ADMET and toxicity leaderboards

TDC ADMET: Pgp_Broccatelli

Official TDC ADMET Group leaderboard task for the Pgp_Broccatelli endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity...

ADMET and toxicity leaderboardsTherapeutics Data Commons ADMET Groupchemistry benchmarkleaderboardTDC
Benchmark · ADMET and toxicity leaderboards

TDC ADMET: Solubility_AqSolDB

Official TDC ADMET Group leaderboard task for the Solubility_AqSolDB endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxici...

ADMET and toxicity leaderboardsTherapeutics Data Commons ADMET Groupchemistry benchmarkleaderboardTDC
Benchmark · ADMET and toxicity leaderboards

TDC ADMET: VDss_Lombardo

Official TDC ADMET Group leaderboard task for the VDss_Lombardo endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity le...

ADMET and toxicity leaderboardsTherapeutics Data Commons ADMET Groupchemistry benchmarkleaderboardTDC
Benchmark · ADMET and toxicity leaderboards

TDC ADMET: hERG

Official TDC ADMET Group leaderboard task for the hERG endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity leaderboard...

ADMET and toxicity leaderboardsTherapeutics Data Commons ADMET Groupchemistry benchmarkleaderboardTDC
Benchmark · Perturbation and trial outcome

TDC Clinical Trial Outcome Prediction: TOP

Clinical trial outcome prediction benchmark for modeling trial success probability. Benchmark suite: Therapeutics Data Commons Clinical Trial Group. Primary category: Perturbation and trial outcome. Leaderboard status: active TDC leaderboa...

Perturbation and trial outcomeTherapeutics Data Commons Clinical Trial Groupchemistry benchmarkleaderboardTDC
Benchmark · Perturbation and trial outcome

TDC Counterfactual: scperturb_drug_SrivatsanTrapnell2020_sciplex2

Drug perturbation benchmark for counterfactual response modeling in single-cell experiments. Benchmark suite: Therapeutics Data Commons Counterfactual Group. Primary category: Perturbation and trial outcome. Leaderboard status: active TDC...

Perturbation and trial outcomeTherapeutics Data Commons Counterfactual Groupchemistry benchmarkleaderboardTDC
Benchmark · DTI and binding prediction

TDC DTI-DG: BindingDB_Patent

Domain-generalization benchmark for drug-target interaction prediction from BindingDB patent data. Benchmark suite: Therapeutics Data Commons DTI Domain Generalization Group. Primary category: DTI and binding prediction. Leaderboard status...

DTI and binding predictionTherapeutics Data Commons DTI Domain Generalization Groupchemistry benchmarkleaderboardTDC
Benchmark · Docking and binding benchmarks

TDC Docking: DRD3

TDC docking benchmark around DRD3, useful for comparing docking and pose-scoring workflows. Benchmark suite: Therapeutics Data Commons Docking Group. Primary category: Docking and binding benchmarks. Leaderboard status: active TDC leaderbo...

Docking and binding benchmarksTherapeutics Data Commons Docking Groupchemistry benchmarkleaderboardTDC
Benchmark · Drug combination leaderboards

TDC DrugCombo: DrugComb_Bliss

Official TDC drug-combination response benchmark for DrugComb_Bliss, supporting model comparison on synergy and combination-effect prediction. Benchmark suite: Therapeutics Data Commons DrugCombo Group. Primary category: Drug combination l...

Drug combination leaderboardsTherapeutics Data Commons DrugCombo Groupchemistry benchmarkleaderboardTDC
Benchmark · Drug combination leaderboards

TDC DrugCombo: DrugComb_CSS

Official TDC drug-combination response benchmark for DrugComb_CSS, supporting model comparison on synergy and combination-effect prediction. Benchmark suite: Therapeutics Data Commons DrugCombo Group. Primary category: Drug combination lea...

Drug combination leaderboardsTherapeutics Data Commons DrugCombo Groupchemistry benchmarkleaderboardTDC
Benchmark · Drug combination leaderboards

TDC DrugCombo: DrugComb_HSA

Official TDC drug-combination response benchmark for DrugComb_HSA, supporting model comparison on synergy and combination-effect prediction. Benchmark suite: Therapeutics Data Commons DrugCombo Group. Primary category: Drug combination lea...

Drug combination leaderboardsTherapeutics Data Commons DrugCombo Groupchemistry benchmarkleaderboardTDC
Benchmark · Drug combination leaderboards

TDC DrugCombo: DrugComb_Loewe

Official TDC drug-combination response benchmark for DrugComb_Loewe, supporting model comparison on synergy and combination-effect prediction. Benchmark suite: Therapeutics Data Commons DrugCombo Group. Primary category: Drug combination l...

Drug combination leaderboardsTherapeutics Data Commons DrugCombo Groupchemistry benchmarkleaderboardTDC
Benchmark · Drug combination leaderboards

TDC DrugCombo: DrugComb_ZIP

Official TDC drug-combination response benchmark for DrugComb_ZIP, supporting model comparison on synergy and combination-effect prediction. Benchmark suite: Therapeutics Data Commons DrugCombo Group. Primary category: Drug combination lea...

Drug combination leaderboardsTherapeutics Data Commons DrugCombo Groupchemistry benchmarkleaderboardTDC
Benchmark · Generative molecular design benchmarks

TDC Oracle: DRD2

DRD2 activity oracle for molecular generation. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary category: Gener...

Generative molecular design benchmarksTherapeutics Data Commons Oracleschemistry benchmarkleaderboardTDC
Benchmark · Generative molecular design benchmarks

TDC Oracle: GSK3B

GSK3B activity oracle for molecular generation. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary category: Gene...

Generative molecular design benchmarksTherapeutics Data Commons Oracleschemistry benchmarkleaderboardTDC
Benchmark · Generative molecular design benchmarks

TDC Oracle: JNK3

JNK3 activity oracle for molecular generation. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary category: Gener...

Generative molecular design benchmarksTherapeutics Data Commons Oracleschemistry benchmarkleaderboardTDC
Benchmark · Generative molecular design benchmarks

TDC Oracle: QED

Drug-likeness oracle based on QED. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary category: Generative molecu...

Generative molecular design benchmarksTherapeutics Data Commons Oracleschemistry benchmarkleaderboardTDC
Benchmark · Generative molecular design benchmarks

TDC Oracle: albuterol similarity

Similarity benchmark around albuterol. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary category: Generative mo...

Generative molecular design benchmarksTherapeutics Data Commons Oracleschemistry benchmarkleaderboardTDC
Benchmark · Generative molecular design benchmarks

TDC Oracle: aripiprazole similarity

Similarity benchmark around aripiprazole. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary category: Generative...

Generative molecular design benchmarksTherapeutics Data Commons Oracleschemistry benchmarkleaderboardTDC
Benchmark · Generative molecular design benchmarks

TDC Oracle: celecoxib rediscovery

Rediscovery benchmark targeting celecoxib-like molecules. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary cate...

Generative molecular design benchmarksTherapeutics Data Commons Oracleschemistry benchmarkleaderboardTDC
Benchmark · Generative molecular design benchmarks

TDC Oracle: decorator hop

Decoration-oriented molecular design oracle for scaffold optimization. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles....

Generative molecular design benchmarksTherapeutics Data Commons Oracleschemistry benchmarkleaderboardTDC
Benchmark · Generative molecular design benchmarks

TDC Oracle: fexofenadine MPO

Multi-property optimization benchmark inspired by fexofenadine design. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles....

Generative molecular design benchmarksTherapeutics Data Commons Oracleschemistry benchmarkleaderboardTDC
Benchmark · Generative molecular design benchmarks

TDC Oracle: isomers C7H8N2O2

Formula-constrained isomer generation benchmark for C7H8N2O2. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary...

Generative molecular design benchmarksTherapeutics Data Commons Oracleschemistry benchmarkleaderboardTDC
Benchmark · Generative molecular design benchmarks

TDC Oracle: isomers C9H10N2O2PF2Cl

Formula-constrained isomer generation benchmark for C9H10N2O2PF2Cl. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Pr...

Generative molecular design benchmarksTherapeutics Data Commons Oracleschemistry benchmarkleaderboardTDC
Benchmark · Generative molecular design benchmarks

TDC Oracle: median molecules 1

Median-molecule generation benchmark used in GuacaMol-style evaluations. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracle...

Generative molecular design benchmarksTherapeutics Data Commons Oracleschemistry benchmarkleaderboardTDC
Benchmark · Generative molecular design benchmarks

TDC Oracle: median molecules 2

Second median-molecule generation benchmark for distributional molecule design. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons...

Generative molecular design benchmarksTherapeutics Data Commons Oracleschemistry benchmarkleaderboardTDC
Benchmark · Generative molecular design benchmarks

TDC Oracle: mestranol similarity

Similarity benchmark around mestranol. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary category: Generative mo...

Generative molecular design benchmarksTherapeutics Data Commons Oracleschemistry benchmarkleaderboardTDC
Benchmark · Generative molecular design benchmarks

TDC Oracle: osimertinib MPO

Multi-property optimization benchmark inspired by osimertinib design. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles....

Generative molecular design benchmarksTherapeutics Data Commons Oracleschemistry benchmarkleaderboardTDC
Benchmark · Generative molecular design benchmarks

TDC Oracle: penalized logP

Penalized logP optimization oracle for de novo design. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary categor...

Generative molecular design benchmarksTherapeutics Data Commons Oracleschemistry benchmarkleaderboardTDC
Benchmark · Generative molecular design benchmarks

TDC Oracle: scaffold hop

Scaffold hopping oracle for generative molecular design. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary categ...

Generative molecular design benchmarksTherapeutics Data Commons Oracleschemistry benchmarkleaderboardTDC
Benchmark · Generative molecular design benchmarks

TDC Oracle: synthetic accessibility

Synthetic-accessibility oracle for generative molecule design. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary...

Generative molecular design benchmarksTherapeutics Data Commons Oracleschemistry benchmarkleaderboardTDC
Benchmark · Generative molecular design benchmarks

TDC Oracle: thiothixene rediscovery

Rediscovery benchmark targeting thiothixene-like molecules. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary ca...

Generative molecular design benchmarksTherapeutics Data Commons Oracleschemistry benchmarkleaderboardTDC
Benchmark · Generative molecular design benchmarks

TDC Oracle: troglitazone rediscovery

Rediscovery benchmark targeting troglitazone-like molecules. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary c...

Generative molecular design benchmarksTherapeutics Data Commons Oracleschemistry benchmarkleaderboardTDC
Benchmark · Generative molecular design benchmarks

TDC Oracle: valsartan SMARTS

SMARTS-constrained molecular generation benchmark around valsartan-like chemistry. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Comm...

Generative molecular design benchmarksTherapeutics Data Commons Oracleschemistry benchmarkleaderboardTDC
Benchmark · Peptide and protein binding benchmarks

TDC Protein-Peptide: brown_mdm2_ace2_12ca5

Protein-peptide benchmark for peptide binding and design model evaluation. Benchmark suite: Therapeutics Data Commons Protein-Peptide Group. Primary category: Peptide and protein binding benchmarks. Leaderboard status: active TDC leaderboa...

Peptide and protein binding benchmarksTherapeutics Data Commons Protein-Peptide Groupchemistry benchmarkleaderboardTDC
Benchmark · Peptide and protein binding benchmarks

TDC Protein-Peptide: tchard

Protein-peptide benchmark for evaluating peptide binding predictors. Benchmark suite: Therapeutics Data Commons Protein-Peptide Group. Primary category: Peptide and protein binding benchmarks. Leaderboard status: active TDC leaderboard.

Peptide and protein binding benchmarksTherapeutics Data Commons Protein-Peptide Groupchemistry benchmarkleaderboardTDC
Benchmark · DTI and binding prediction

TDC Single-cell DTI: opentargets_dti

Single-cell drug-target interaction benchmark built around Open Targets data. Benchmark suite: Therapeutics Data Commons Single-cell DTI Group. Primary category: DTI and binding prediction. Leaderboard status: active TDC leaderboard.

DTI and binding predictionTherapeutics Data Commons Single-cell DTI Groupchemistry benchmarkleaderboardTDC
Benchmark · DFT workflows, Generative materials design

Matbench Discovery:benchmarking machine learning energy models for materials discovery.

Matbench Discovery is a platform for benchmarking machine learning energy models for materials discovery. It provides a comprehensive test set and metrics to evaluate model performance, including metrics like CPS (Composite Performance Sco...

materials sciencemachine learningbenchmarkingenergy modelsmaterials discovery
Benchmark · Characterization & Anallysis

MatQnA: A Benchmark Dataset for Multi-modal LLMs in Materials Characterization

MatQnA is the first multi-modal benchmark dataset designed for evaluating large language models (LLMs) in materials characterization and analysis. It covers ten mainstream characterization methods and includes both multiple-choice and subj...

materials sciencebenchmark datasetmulti-modal LLMscharacterizationXPS
Benchmark

MatSciBench: Materials Science Reasoning Benchmark

MatSciBench is a college-level benchmark designed to evaluate the reasoning capabilities of large language models in materials science. It includes 1,340 problems covering quantitative, symbolic, and multimodal question answering, with ref...

materials sciencebenchmarklarge language modelsreasoningquantitative analysis
Benchmark

MaCBench: Multimodal Reasoning Benchmark for Chemistry and Materials Science

MaCBench is a benchmark designed to evaluate the multimodal reasoning capabilities of vision-language models (VLMs) in chemistry and materials science. It includes over 1,100 hand-crafted question-image pairs covering data extraction, expe...

multimodal reasoningchemistrymaterials sciencebenchmarkvision-language models
Benchmark

JARVIS-Leaderboard

This project benchmarks the performance of materials-science methods using datasets from the JARVIS-Tools databases. It covers AI-driven materials design, property prediction, electronic-structure methods, force fields, quantum computing,...

materials sciencebenchmarkAImachine learningJARVIS
Benchmark · Generative materials design

LeMat-GenBench: Generative Crystal Structure Model Benchmark

A Hugging Face Space that provides a leaderboard for ranking generative models of crystal structures based on various quality metrics. Users can upload their own CIF files and model details to be included in the benchmark.

crystal structuregenerative modelsbenchmarkmaterials scienceHugging Face