CASF-2016 ranking benchmark
CASF benchmark task for ranking protein-ligand complexes by affinity. Benchmark suite: CASF. Primary category: Docking and binding benchmarks. Leaderboard status: benchmark suite with published rankings.
Benchmark pages help researchers compare models on shared tasks, reusable evaluation protocols, held-out datasets, and community challenge results.
Chemistry benchmarks make model claims easier to compare by tying tasks to datasets, metrics, splits, and evaluation protocols. This hub collects benchmark activities across molecular property prediction, generative molecule design, retrosynthesis, reaction prediction, spectra interpretation, docking, and virtual screening workflows.
107 approved benchmark entries are available for search indexing and community discovery.
CASF benchmark task for ranking protein-ligand complexes by affinity. Benchmark suite: CASF. Primary category: Docking and binding benchmarks. Leaderboard status: benchmark suite with published rankings.
Comparative Assessment of Scoring Functions benchmark for protein-ligand scoring power. Benchmark suite: CASF. Primary category: Docking and binding benchmarks. Leaderboard status: benchmark suite with published rankings.
Blind benchmark for ranking protein-ligand binding affinities across challenge targets. Benchmark suite: D3R Grand Challenges. Primary category: Docking and binding benchmarks. Leaderboard status: community blind challenge with public resu...
Blind challenge benchmark for binding free-energy prediction. Benchmark suite: D3R Grand Challenges. Primary category: Docking and binding benchmarks. Leaderboard status: community blind challenge with public results.
Blind docking benchmark focused on predicting protein-ligand binding poses. Benchmark suite: D3R Grand Challenges. Primary category: Docking and binding benchmarks. Leaderboard status: community blind challenge with public results.
Directory of Useful Decoys Enhanced benchmark for structure-based virtual screening. Benchmark suite: DUD-E. Primary category: Virtual screening benchmarks. Leaderboard status: benchmark dataset with extensive public comparisons.
Benchmark suite for goal-directed and distribution-learning molecular generation. Benchmark suite: GuacaMol. Primary category: Generative molecular design benchmarks. Leaderboard status: published benchmark suite with reference leaderboard...
NIST JARVIS leaderboard collection for AI models predicting material and molecular properties. Benchmark suite: JARVIS-Leaderboard. Primary category: Materials and spectra benchmarks. Leaderboard status: public JARVIS leaderboard.
JARVIS atomistic generation benchmark tasks for structure generation models. Benchmark suite: JARVIS-Leaderboard. Primary category: Generative materials benchmarks. Leaderboard status: public JARVIS leaderboard.
JARVIS benchmark family for machine-learning force fields and atomistic simulations. Benchmark suite: JARVIS-Leaderboard. Primary category: Catalysis and atomistic simulation. Leaderboard status: public JARVIS leaderboard.
JARVIS electronic-structure benchmarks for band gaps, energetics, and related properties. Benchmark suite: JARVIS-Leaderboard. Primary category: Quantum chemistry benchmarks. Leaderboard status: public JARVIS leaderboard.
JARVIS benchmarks for spectral-property prediction and analysis. Benchmark suite: JARVIS-Leaderboard. Primary category: Spectra and analytical benchmarks. Leaderboard status: public JARVIS leaderboard.
Literature-derived PubChem BioAssay benchmark for virtual screening and target-specific ranking. Benchmark suite: LIT-PCBA. Primary category: Virtual screening benchmarks. Leaderboard status: benchmark dataset with public method comparison...
Molecular generation benchmark platform with datasets, metrics, and baseline generative models. Benchmark suite: MOSES. Primary category: Generative molecular design benchmarks. Leaderboard status: benchmark platform with baseline comparis...
Benchmark for materials discovery and stability prediction models. Benchmark suite: Matbench Discovery. Primary category: Materials and molecular property benchmarks. Leaderboard status: public benchmark suite and leaderboard.
Materials machine-learning benchmark suite, useful for ChemistryAtlas users working across molecular and materials discovery. Benchmark suite: Matbench. Primary category: Materials and molecular property benchmarks. Leaderboard status: pub...
MoleculeNet benchmark task for BACE inhibitor classification. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Bioactivity prediction bench...
MoleculeNet benchmark task for blood-brain barrier penetration classification. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: ADMET and t...
MoleculeNet benchmark task for organic photovoltaic candidate property prediction. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Materia...
MoleculeNet benchmark task for clinical trial toxicity classification. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: ADMET and toxicity...
MoleculeNet benchmark task for aqueous solubility regression. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Physical chemistry benchmark...
MoleculeNet benchmark task for hydration free energy regression. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Physical chemistry benchm...
MoleculeNet benchmark task for anti-HIV activity classification. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Bioactivity prediction be...
MoleculeNet benchmark task for octanol-water distribution regression. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Physical chemistry b...
MoleculeNet benchmark task for unbiased virtual-screening benchmark. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Virtual screening ben...
MoleculeNet benchmark task for PubChem BioAssay multi-task benchmark. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Bioactivity predicti...
MoleculeNet benchmark task for protein-ligand binding affinity prediction. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Docking and bin...
MoleculeNet benchmark task for quantum property regression. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Quantum chemistry benchmarks....
MoleculeNet benchmark task for multi-property quantum chemistry regression. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Quantum chemis...
MoleculeNet benchmark task for excited-state quantum-property prediction. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Quantum chemistr...
MoleculeNet benchmark task for small-molecule quantum-property prediction. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: Quantum chemist...
MoleculeNet benchmark task for drug side-effect classification. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: ADMET and toxicity leaderb...
MoleculeNet benchmark task for multi-task toxicity classification. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: ADMET and toxicity lead...
MoleculeNet benchmark task for high-throughput toxicity assay classification. It is widely used to compare molecular graph, fingerprint, and foundation-model property predictors. Benchmark suite: MoleculeNet. Primary category: ADMET and to...
Large-scale molecular graph benchmark for HOMO-LUMO gap prediction from quantum chemistry labels. Benchmark suite: Open Graph Benchmark Large-Scale Challenge. Primary category: Molecular graph and property prediction. Leaderboard status: p...
Initial-structure-to-relaxed-energy task for catalyst adsorbate systems. Benchmark suite: Open Catalyst Project. Primary category: Catalysis and atomistic simulation. Leaderboard status: public Open Catalyst leaderboard.
Initial-structure-to-relaxed-structure benchmark for catalyst systems. Benchmark suite: Open Catalyst Project. Primary category: Catalysis and atomistic simulation. Leaderboard status: public Open Catalyst leaderboard.
Structure-to-energy-and-forces benchmark for adsorbate-catalyst systems. Benchmark suite: Open Catalyst Project. Primary category: Catalysis and atomistic simulation. Leaderboard status: public Open Catalyst leaderboard.
Open Catalyst 2022 benchmark for oxide catalyst systems and atomistic ML. Benchmark suite: Open Catalyst Project. Primary category: Catalysis and atomistic simulation. Leaderboard status: Open Catalyst benchmark and challenge ecosystem.
Emerging Open Catalyst benchmark set for next-generation catalyst modeling. Benchmark suite: Open Catalyst Project. Primary category: Catalysis and atomistic simulation. Leaderboard status: emerging Open Catalyst benchmark.
Benchmark and live leaderboard for Raman spectroscopy AI models. Benchmark suite: RamanBench. Primary category: Spectra and analytical benchmarks. Leaderboard status: published benchmark with live leaderboard.
Blind host-guest binding free-energy challenge for force-field, simulation, and ML methods. Benchmark suite: SAMPL Challenges. Primary category: Docking and binding benchmarks. Leaderboard status: community blind challenge with public resu...
Blind challenge benchmark for hydration free-energy prediction using physical and ML methods. Benchmark suite: SAMPL Challenges. Primary category: Physical chemistry benchmarks. Leaderboard status: community blind challenge with public res...
Blind partition/distribution coefficient challenge for logP and logD prediction. Benchmark suite: SAMPL Challenges. Primary category: Physical chemistry benchmarks. Leaderboard status: community blind challenge with public results.
Blind pKa prediction challenge for small-molecule protonation and thermodynamics methods. Benchmark suite: SAMPL Challenges. Primary category: Physical chemistry benchmarks. Leaderboard status: community blind challenge with public results.
Official TDC ADMET Group leaderboard task for the AMES endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity leaderboard...
Official TDC ADMET Group leaderboard task for the BBB_Martins endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity lead...
Official TDC ADMET Group leaderboard task for the Bioavailability_Ma endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxici...
Official TDC ADMET Group leaderboard task for the CYP2C9_Substrate_CarbonMangels endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADME...
Official TDC ADMET Group leaderboard task for the CYP2C9_Veith endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity lea...
Official TDC ADMET Group leaderboard task for the CYP2D6_Substrate_CarbonMangels endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADME...
Official TDC ADMET Group leaderboard task for the CYP2D6_Veith endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity lea...
Official TDC ADMET Group leaderboard task for the CYP3A4_Substrate_CarbonMangels endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADME...
Official TDC ADMET Group leaderboard task for the CYP3A4_Veith endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity lea...
Official TDC ADMET Group leaderboard task for the Caco2_Wang endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity leade...
Official TDC ADMET Group leaderboard task for the Clearance_Hepatocyte_AZ endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and t...
Official TDC ADMET Group leaderboard task for the Clearance_Microsome_AZ endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and to...
Official TDC ADMET Group leaderboard task for the DILI endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity leaderboard...
Official TDC ADMET Group leaderboard task for the HIA_Hou endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity leaderbo...
Official TDC ADMET Group leaderboard task for the Half_Life_Obach endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity...
Official TDC ADMET Group leaderboard task for the LD50_Zhu endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity leaderb...
Official TDC ADMET Group leaderboard task for the Lipophilicity_AstraZeneca endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and...
Official TDC ADMET Group leaderboard task for the PPBR_AZ endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity leaderbo...
Official TDC ADMET Group leaderboard task for the Pgp_Broccatelli endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity...
Official TDC ADMET Group leaderboard task for the Solubility_AqSolDB endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxici...
Official TDC ADMET Group leaderboard task for the VDss_Lombardo endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity le...
Official TDC ADMET Group leaderboard task for the hERG endpoint, useful for benchmarking small-molecule ADMET and toxicity predictors. Benchmark suite: Therapeutics Data Commons ADMET Group. Primary category: ADMET and toxicity leaderboard...
Clinical trial outcome prediction benchmark for modeling trial success probability. Benchmark suite: Therapeutics Data Commons Clinical Trial Group. Primary category: Perturbation and trial outcome. Leaderboard status: active TDC leaderboa...
Drug perturbation benchmark for counterfactual response modeling in single-cell experiments. Benchmark suite: Therapeutics Data Commons Counterfactual Group. Primary category: Perturbation and trial outcome. Leaderboard status: active TDC...
Domain-generalization benchmark for drug-target interaction prediction from BindingDB patent data. Benchmark suite: Therapeutics Data Commons DTI Domain Generalization Group. Primary category: DTI and binding prediction. Leaderboard status...
TDC docking benchmark around DRD3, useful for comparing docking and pose-scoring workflows. Benchmark suite: Therapeutics Data Commons Docking Group. Primary category: Docking and binding benchmarks. Leaderboard status: active TDC leaderbo...
Official TDC drug-combination response benchmark for DrugComb_Bliss, supporting model comparison on synergy and combination-effect prediction. Benchmark suite: Therapeutics Data Commons DrugCombo Group. Primary category: Drug combination l...
Official TDC drug-combination response benchmark for DrugComb_CSS, supporting model comparison on synergy and combination-effect prediction. Benchmark suite: Therapeutics Data Commons DrugCombo Group. Primary category: Drug combination lea...
Official TDC drug-combination response benchmark for DrugComb_HSA, supporting model comparison on synergy and combination-effect prediction. Benchmark suite: Therapeutics Data Commons DrugCombo Group. Primary category: Drug combination lea...
Official TDC drug-combination response benchmark for DrugComb_Loewe, supporting model comparison on synergy and combination-effect prediction. Benchmark suite: Therapeutics Data Commons DrugCombo Group. Primary category: Drug combination l...
Official TDC drug-combination response benchmark for DrugComb_ZIP, supporting model comparison on synergy and combination-effect prediction. Benchmark suite: Therapeutics Data Commons DrugCombo Group. Primary category: Drug combination lea...
DRD2 activity oracle for molecular generation. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary category: Gener...
GSK3B activity oracle for molecular generation. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary category: Gene...
JNK3 activity oracle for molecular generation. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary category: Gener...
Drug-likeness oracle based on QED. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary category: Generative molecu...
Similarity benchmark around albuterol. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary category: Generative mo...
Similarity benchmark around aripiprazole. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary category: Generative...
Rediscovery benchmark targeting celecoxib-like molecules. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary cate...
Decoration-oriented molecular design oracle for scaffold optimization. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles....
Multi-property optimization benchmark inspired by fexofenadine design. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles....
Formula-constrained isomer generation benchmark for C7H8N2O2. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary...
Formula-constrained isomer generation benchmark for C9H10N2O2PF2Cl. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Pr...
Median-molecule generation benchmark used in GuacaMol-style evaluations. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracle...
Second median-molecule generation benchmark for distributional molecule design. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons...
Similarity benchmark around mestranol. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary category: Generative mo...
Multi-property optimization benchmark inspired by osimertinib design. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles....
Penalized logP optimization oracle for de novo design. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary categor...
Scaffold hopping oracle for generative molecular design. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary categ...
Synthetic-accessibility oracle for generative molecule design. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary...
Rediscovery benchmark targeting thiothixene-like molecules. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary ca...
Rediscovery benchmark targeting troglitazone-like molecules. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Commons Oracles. Primary c...
SMARTS-constrained molecular generation benchmark around valsartan-like chemistry. TDC oracles are useful for testing generative model optimization behavior, validity, novelty, and design objectives. Benchmark suite: Therapeutics Data Comm...
Protein-peptide benchmark for peptide binding and design model evaluation. Benchmark suite: Therapeutics Data Commons Protein-Peptide Group. Primary category: Peptide and protein binding benchmarks. Leaderboard status: active TDC leaderboa...
Protein-peptide benchmark for evaluating peptide binding predictors. Benchmark suite: Therapeutics Data Commons Protein-Peptide Group. Primary category: Peptide and protein binding benchmarks. Leaderboard status: active TDC leaderboard.
Single-cell drug-target interaction benchmark built around Open Targets data. Benchmark suite: Therapeutics Data Commons Single-cell DTI Group. Primary category: DTI and binding prediction. Leaderboard status: active TDC leaderboard.
MatBench is an automated leaderboard for benchmarking state-of-the-art machine learning algorithms for predicting a diverse range of solid materials' properties. It is hosted and maintained by the Materials Project. Problems Matbench v0.1...
Matbench Discovery is a platform for benchmarking machine learning energy models for materials discovery. It provides a comprehensive test set and metrics to evaluate model performance, including metrics like CPS (Composite Performance Sco...
MatQnA is the first multi-modal benchmark dataset designed for evaluating large language models (LLMs) in materials characterization and analysis. It covers ten mainstream characterization methods and includes both multiple-choice and subj...
MatSciBench is a college-level benchmark designed to evaluate the reasoning capabilities of large language models in materials science. It includes 1,340 problems covering quantitative, symbolic, and multimodal question answering, with ref...
MaCBench is a benchmark designed to evaluate the multimodal reasoning capabilities of vision-language models (VLMs) in chemistry and materials science. It includes over 1,100 hand-crafted question-image pairs covering data extraction, expe...
This project benchmarks the performance of materials-science methods using datasets from the JARVIS-Tools databases. It covers AI-driven materials design, property prediction, electronic-structure methods, force fields, quantum computing,...
A Hugging Face Space that provides a leaderboard for ranking generative models of crystal structures based on various quality metrics. Users can upload their own CIF files and model details to be included in the benchmark.