{"categories":[{"id":"platform","name":"Platforms","blurb":"Umbrella initiatives that host many benchmarks, define evaluation protocols, and run leaderboards. Start here if you want one dependency instead of twelve."},{"id":"admet","name":"Molecular Property & ADMET Prediction","blurb":"Absorption, distribution, metabolism, excretion, toxicity, solubility, permeability. The most crowded and most criticised corner of the field — beware of leaky splits and duplicate assay records."},{"id":"docking","name":"Docking, Pose Prediction & Co-folding","blurb":"Can your model put the ligand in the right place, in a physically legal conformation? Physical-plausibility checks and time-split holdouts matter more than raw RMSD here."},{"id":"affinity","name":"Binding Affinity & Free Energy","blurb":"Scoring functions, relative and absolute binding free energy, and the datasets they are calibrated on. Watch for train/test similarity between PDBbind-derived sets."},{"id":"generative","name":"Generative & De Novo Design","blurb":"Distribution learning, goal-directed optimisation, and structure-based generation. Sample-efficiency-aware benchmarks have largely superseded the unconstrained oracle-call benchmarks."},{"id":"structure","name":"Protein & Complex Structure Prediction","blurb":"Blind community assessments and fitness-landscape benchmarks for structure, function and variant effect prediction."},{"id":"synthesis","name":"Retrosynthesis & Synthesis Planning","blurb":"Single-step retrosynthesis, multi-step route planning, reaction outcome and condition prediction. Route-level evaluation is much harder than top-k single-step accuracy."},{"id":"target","name":"Target ID, Omics & Perturbation","blurb":"Target identification, knowledge graphs, perturbation response prediction and single-cell tasks that feed the earliest stage of the pipeline."},{"id":"clinical","name":"Clinical & Translational","blurb":"Trial outcome prediction, patient-level and real-world-evidence tasks. Small n, heavy confounding, and the highest stakes."},{"id":"agent","name":"LLM & Agent Benchmarks","blurb":"Evaluations of language models and autonomous agents on chemistry, biology and end-to-end discovery workflows. The fastest-moving category — expect churn."}],"consortia":{"TDC":"Therapeutics Data Commons — an open-science initiative led by the Zitnik Lab (Harvard Medical School, MIMS) with contributors from Stanford, MIT, UIUC and IQVIA. Coordinates AI-ready datasets, oracles and leaderboards across the whole therapeutic pipeline. Community runs on Slack and a public mailing list.","Polaris":"The Polaris consortium — a cross-industry benchmarking effort convened by Valence Labs / Recursion with representatives from AstraZeneca, Merck, Novartis, Johnson & Johnson, Pfizer, Bayer, Relay Therapeutics, Nimbus Therapeutics and Blueprint Medicines. Its emphasis is on curation standards and domain-appropriate evaluation rather than on maximising dataset count.","OMSF":"Open Molecular Software Foundation — a non-profit fiscal and governance host for open computational-chemistry projects including Open Force Field, OpenFE and OpenADMET. Provides neutral stewardship so benchmarks are not owned by a single company or lab.","OpenFF":"Open Force Field Initiative / Consortium — academic and industry partners (including most large pharma) building open, reproducible force fields and the benchmark sets used to validate them. Hosted under OMSF.","ASAP":"AI-driven Structure-enabled Antiviral Platform — an open-science antiviral discovery consortium funded under the NIH/NIAID AViDD programme, spanning UCSF, Diamond Light Source, the University of Oxford, MSKCC and others. Releases all data openly and runs blinded prediction challenges through Polaris.","SGC":"Structural Genomics Consortium — a long-running public-private partnership (pharma + academia + charitable funders) with a strict open-science, no-patent policy. Runs the CACHE blinded hit-finding challenges.","MLPDS":"Machine Learning for Pharmaceutical Discovery and Synthesis Consortium — MIT-led industry consortium (multiple large pharma members) behind ASKCOS and much of the open retrosynthesis tooling.","OpenTargets":"Open Targets — a public-private partnership between EMBL-EBI, the Wellcome Sanger Institute, GSK, Sanofi, Genentech, Pfizer and MSD, focused on systematic target identification and prioritisation.","DeepChem":"DeepChem — a community open-source project (originating at Stanford's Pande Lab) that maintains MoleculeNet and a broad library of chemistry ML tooling.","OpenProblems":"Open Problems in Single-Cell Analysis — a CZI-supported community consortium that maintains continuously re-run, formalised benchmarks for single-cell tasks.","FutureHouse":"FutureHouse — a non-profit research lab building AI scientists, which releases open evaluation suites for literature research and bioinformatics agents.","D3R":"Drug Design Data Resource — an NIH-funded effort at UC San Diego that ran the blinded Grand Challenge series on pose prediction, affinity ranking and free energy.","CASP":"Critical Assessment of Structure Prediction — the long-running biennial blinded community experiment coordinated by the Protein Structure Prediction Center (UC Davis), now including protein-ligand complex categories.","PLINDER":"PLINDER / OpenBind-adjacent effort — a collaboration between Exscientia, NVIDIA and academic groups producing large, leakage-controlled protein-ligand interaction datasets and co-folding benchmarks.","None":"","Insilico":"Insilico Medicine — a clinical-stage generative AI drug discovery company (HKEX: 3696). Its MMAI Gym programme releases benchmarks and public leaderboards for evaluating frontier and foundation models on real drug discovery tasks. The distinguishing claim is decontamination: benchmarks are anchored in internally validated programmes and held-out real-world data rather than public sets that may already sit in model training corpora."},"benchmarks":[{"name":"Therapeutics Data Commons (TDC / TDC-2)","cat":"platform","author":"Kexin Huang & Marinka Zitnik","repo":"mims-harvard/TDC","site":"https://tdcommons.ai","consortium":"TDC","desc":"The broadest single entry point in the field: 60+ AI-ready datasets organised as single-instance prediction, multi-instance prediction and generation, with prescribed splits, evaluators, generation oracles and public leaderboards. TDC-2 extends into multimodal and single-cell data plus a model hub.","tasks":"ADMET, DTI, HTS, generation, trial outcome, single-cell","note":"Because so many downstream papers report TDC ADMET numbers, split fidelity matters enormously here — always use the provided splits rather than re-splitting.","added":"2026-07-25"},{"name":"PyTDC","cat":"platform","author":"Alejandro Velez-Arce & Marinka Zitnik","repo":"apliko-xyz/PyTDC","site":"https://pytdc.apliko.io","consortium":"None","desc":"A machine-learning platform for biomedical foundation models, covering training, evaluation and inference rather than datasets alone. Introduces an API-first dataset architecture that unifies heterogeneous, continuously updated sources, and a model server giving standardised inference endpoints and unified access to model weights across distributed repositories. Adds contextualised single-cell therapeutics tasks. ICML 2025.","tasks":"Foundation model training, evaluation, inference, single-cell tasks","note":"Forked from mims-harvard/TDC and now developed separately, so the two diverge and results are not automatically comparable across them. Note the package name: `pip install pytdc-nextml` installs PyTDC, while `pip install PyTDC` installs the original TDC package — an easy and consequential mix-up. TDC is a dataset store; PyTDC is infrastructure built on top of it.","added":"2026-08-17"},{"name":"Polaris","cat":"platform","author":"Cas Wognum & the Polaris team (Valence Labs)","repo":"polaris-hub/polaris","site":"https://polarishub.io","consortium":"Polaris","desc":"A benchmarking platform built around the argument that most public ML-for-drug-discovery benchmarks are too noisy or too easy to be informative. Datasets are expert-curated, versioned and certified; benchmarks bind a dataset to a specific split and metric so results are comparable by construction.","tasks":"Property prediction, potency, ADMET, blinded challenges","note":"Deliberately smaller and stricter than TDC. Blinded challenge results (e.g. the ASAP antiviral challenges) are the most informative part.","added":"2026-07-25"},{"name":"MoleculeNet","cat":"platform","author":"Zhenqin Wu, Bharath Ramsundar & Vijay Pande","repo":"deepchem/deepchem","site":"https://moleculenet.org","consortium":"DeepChem","desc":"The original unified benchmark for molecular machine learning, covering quantum mechanics, physical chemistry, biophysics and physiology tasks (QM9, ESOL, FreeSolv, Lipophilicity, BACE, BBBP, Tox21, ToxCast, SIDER, ClinTox, HIV, MUV, PCBA). Distributed through DeepChem.","tasks":"QM, solubility, toxicity, bioactivity","note":"Still the most-cited benchmark and the most criticised: several tasks are tiny, label noise is high, and random splits leak badly. Prefer scaffold splits and treat small-set results as noise unless averaged over many seeds.","added":"2026-07-25"},{"name":"Open Graph Benchmark (ogbg-mol*)","cat":"platform","author":"Weihua Hu, Matthias Fey & Jure Leskovec","repo":"snap-stanford/ogb","site":"https://ogb.stanford.edu","consortium":"None","desc":"Standardised graph ML benchmarks with fixed scaffold splits and an automated leaderboard. The molecular subsets (ogbg-molhiv, ogbg-molpcba) are derived from MoleculeNet but with far more disciplined evaluation, and ogb-lsc/PCQM4Mv2 is the large-scale quantum property task.","tasks":"Graph property prediction, quantum properties","note":"The gold standard for split hygiene in graph ML, though the chemistry itself is a re-packaging of existing sets.","added":"2026-07-25"},{"name":"OpenADMET","cat":"admet","author":"OpenADMET team (OMSF)","repo":"OpenADMET/openadmet-toolkit","site":"https://openadmet.org","consortium":"OMSF","desc":"An open, pre-competitive effort to generate and release genuinely new ADMET measurements — rather than re-mining ChEMBL — together with open models and evaluation tooling. Ran joint blinded ADMET challenges with ASAP Discovery and Polaris.","tasks":"ADMET, metabolic stability, permeability","note":"One of very few efforts producing new prospective data rather than recycling public assay dumps.","added":"2026-07-25"},{"name":"ASAP Discovery","cat":"platform","author":"ASAP Discovery Consortium","repo":"asapdiscovery/asapdiscovery","site":"https://asapdiscovery.org","consortium":"ASAP","desc":"An open-science antiviral discovery programme that releases structures, assay data and computational tooling openly and in near-real-time, and converts its live campaigns into blinded prediction challenges (potency, ADMET, pose) hosted on Polaris.","tasks":"Potency, ADMET, pose prediction, blinded challenges","note":"The blinded, prospective format makes these among the most trustworthy evaluations available.","added":"2026-07-25"},{"name":"TDC ADMET Benchmark Group","cat":"admet","author":"Kexin Huang & Marinka Zitnik","repo":"mims-harvard/TDC","site":"https://tdcommons.ai/benchmark/admet_group/overview/","consortium":"TDC","desc":"22 ADMET endpoints (Caco-2, HIA, Pgp, bioavailability, Lipophilicity, solubility, PPBR, VDss, CYP isoforms, half-life, clearance, hERG, Ames, DILI, LD50) with scaffold splits, five prescribed seeds and a live leaderboard. The de facto standard reporting format for ADMET papers.","tasks":"22 ADMET endpoints","note":"Several endpoints have only a few hundred compounds, so leaderboard gaps of a few percent are usually within seed noise. Report the standard deviation.","added":"2026-07-25"},{"name":"Biogen ADME Public Dataset","cat":"admet","author":"Cheng Fang et al. (Biogen)","repo":"molecularinformatics/Computational-ADME","site":"","consortium":"None","desc":"Six ADME endpoints (human and rat liver microsomal clearance, MDCK permeability, solubility, human plasma protein binding) measured in a single lab under consistent protocols over several years, released with a prospective time-split.","tasks":"Clearance, permeability, solubility, PPB","note":"Unusually valuable because the assays are internally consistent and the time-split simulates real prospective use. Much harder than public aggregated data suggests.","added":"2026-07-25"},{"name":"ADMET-AI","cat":"admet","author":"Kyle Swanson (Stanford)","repo":"swansonk14/admet_ai","site":"https://admet.ai.greenstonebio.com","consortium":"None","desc":"A fast Chemprop-RDKit model suite trained across the full TDC ADMET collection, with a web interface and the ability to contextualise a prediction against the distribution of approved drugs (DrugBank percentiles).","tasks":"ADMET prediction + reference model","note":"Widely used as the baseline to beat; also a good sanity check on whether a new architecture is actually adding anything.","added":"2026-07-25"},{"name":"Tox21 / DeepTox","cat":"admet","author":"Ruili Huang (NCATS); Andreas Mayr & Sepp Hochreiter (DeepTox)","repo":"","site":"https://tripod.nih.gov/tox21/challenge/","consortium":"None","desc":"The 2014 NIH/NCATS/EPA/FDA challenge measuring ~12,000 compounds against 12 nuclear receptor and stress response pathway assays. The winning DeepTox entry is a landmark result for deep learning in toxicology and the data remains a standard toxicity benchmark.","tasks":"12 toxicity assay endpoints","note":"The original leaderboard is frozen; recent work (2025-26) has proposed reproducible re-implementations because the historical evaluation is hard to replicate exactly.","added":"2026-07-25"},{"name":"MolData","cat":"admet","author":"Arash Keshavarzi Arshadi et al.","repo":"LumosBio/MolData","site":"","consortium":"None","desc":"A large-scale molecular bioactivity benchmark curated from PubChem BioAssay covering ~1.4M compounds across ~600 assay tasks, positioned as a bigger and more realistic alternative to MoleculeNet's PCBA.","tasks":"Multi-task bioactivity classification","note":"Class imbalance is extreme; PR-AUC is far more informative than ROC-AUC here.","added":"2026-07-25"},{"name":"Lo-Hi Benchmark","cat":"admet","author":"Simon Steshin","repo":"SteshinSS/lohi_neurips2023","site":"","consortium":"None","desc":"Splits molecular property tasks into two realistic drug-discovery regimes: Lead Optimisation (Lo, ranking very similar analogues) and Hit Identification (Hi, generalising to dissimilar chemotypes), arguing that conventional benchmarks measure neither.","tasks":"Lead optimisation, hit identification","note":"A useful corrective — many models that look strong on scaffold splits collapse on the Hi setting.","added":"2026-07-25"},{"name":"PoseBusters","cat":"docking","author":"Martin Buttenschoen, Garrett M. Morris & Charlotte M. Deane","repo":"maabuu/posebusters","site":"","consortium":"None","desc":"Both a 428-complex test set of recent, high-quality PDB structures and — more importantly — a validation suite of physical and chemical plausibility checks (bond lengths and angles, stereochemistry, planarity, internal energy, steric clashes with the protein). Redefined how docking success is reported.","tasks":"Pose accuracy + physical validity","note":"The paper's central finding — that several deep-learning docking methods beat classical docking on RMSD while producing physically invalid poses — reshaped the field. 'PB-valid RMSD ≤ 2 Å' is now the expected metric.","added":"2026-07-25"},{"name":"PoseCheck","cat":"docking","author":"Charles Harris (University of Cambridge)","repo":"cch1999/posecheck","site":"","consortium":"None","desc":"An evaluation suite for structure-based generative models that measures steric clashes, ligand strain energy and the profile of recovered protein-ligand interactions, rather than only docking scores.","tasks":"SBDD generative model evaluation","note":"Showed that many 3D generative models produce ligands with implausibly high strain energy — pair it with any SBDD benchmark.","added":"2026-07-25"},{"name":"PoseBench","cat":"docking","author":"Alex Morehead & Jianlin Cheng (University of Missouri)","repo":"BioinfoMachineLearning/PoseBench","site":"","consortium":"None","desc":"A unified harness for comparing protein-ligand structure prediction methods across PoseBusters, Astex Diverse, DockGen and CASP15/16-style ligand sets, including single- and multi-ligand and apo/predicted-structure settings.","tasks":"Cross-method docking comparison","note":"Valuable because it runs the methods rather than trusting reported numbers.","added":"2026-07-25"},{"name":"DockGen","cat":"docking","author":"Gabriele Corso et al. (MIT)","repo":"gcorso/DiffDock","site":"","consortium":"None","desc":"A generalisation benchmark built by clustering binding domains so that test pockets are structurally unlike anything in training. Designed to expose the extent to which docking models memorise pockets rather than learn physics.","tasks":"Blind docking generalisation","note":"Performance drops sharply relative to PoseBusters for most methods — that gap is the point.","added":"2026-07-25"},{"name":"PLINDER","cat":"docking","author":"Janani Durairaj, Yusuf Adeshina et al.","repo":"plinder-org/plinder","site":"https://www.plinder.sh","consortium":"PLINDER","desc":"The largest annotated protein-ligand interaction dataset assembled for ML, with systematic similarity-based splitting to control leakage across sequence, pocket, ligand and interaction space, plus rich per-system quality annotations.","tasks":"Docking / co-folding training + evaluation","note":"The leakage-aware splits are the main contribution; results on random PDBbind splits are not comparable to these.","added":"2026-07-25"},{"name":"Runs N' Poses","cat":"docking","author":"PLINDER team","repo":"plinder-org/runs-n-poses","site":"","consortium":"PLINDER","desc":"A co-folding benchmark for the AlphaFold3-era methods, holding out recent structures and reporting predictions from multiple co-folding models together with PoseBusters physical-validity results for each.","tasks":"Protein-ligand co-folding","note":"Explicitly targets the co-folding setting where the protein structure is predicted rather than given.","added":"2026-07-25"},{"name":"CASF-2016 (Comparative Assessment of Scoring Functions)","cat":"affinity","author":"Minyi Su & Renxiao Wang","repo":"","site":"http://www.pdbbind.org.cn/casf.php","consortium":"None","desc":"The standard four-way assessment of scoring functions: scoring power (correlation with affinity), ranking power (ordering ligands for one target), docking power (identifying the native pose) and screening power (enrichment). Built on a 285-complex core set from PDBbind.","tasks":"Scoring function assessment","note":"Ageing, and heavily overlapped with PDBbind training data — a strong CASF score from a model trained on PDBbind means very little without a leakage analysis.","added":"2026-07-25"},{"name":"PDBbind","cat":"affinity","author":"Renxiao Wang et al.","repo":"","site":"http://www.pdbbind.org.cn","consortium":"None","desc":"The curated collection of PDB complexes with experimentally measured binding affinities (Kd/Ki/IC50) that underpins almost every ML scoring function. Organised into general, refined and core sets.","tasks":"Binding affinity regression","note":"The single biggest source of leakage in the field. Licensing changed to a commercial model for recent versions; several open re-derivations (e.g. via PLINDER, BindingNet) now exist.","added":"2026-07-25"},{"name":"LIT-PCBA","cat":"affinity","author":"Viet-Khoa Tran-Nguyen, Célien Jacquemard & Didier Rognan","repo":"","site":"https://drugdesign.unistra.fr/LIT-PCBA/","consortium":"None","desc":"A virtual screening benchmark drawn from confirmatory PubChem dose-response assays across 15 targets, with unbiased actives and decoys chosen to remove the analogue bias and artificial enrichment that plague DUD-E.","tasks":"Virtual screening enrichment","note":"Much harder and much fairer than DUD-E. Realistic hit rates mean early-enrichment metrics (EF1%, BEDROC) are the ones that matter.","added":"2026-07-25"},{"name":"DUD-E","cat":"affinity","author":"Michael M. Mysinger & Brian K. Shoichet (UCSF)","repo":"","site":"https://dude.docking.org","consortium":"None","desc":"102 targets with actives and property-matched, topologically dissimilar decoys. For a decade the default virtual screening benchmark.","tasks":"Virtual screening enrichment","note":"Now known to contain a strong hidden bias: ML models can separate actives from decoys using ligand features alone, without the protein. Do not use it to validate a structure-based method without a ligand-only control.","added":"2026-07-25"},{"name":"Schrödinger Public Binding Free Energy Benchmark","cat":"affinity","author":"Schrödinger (FEP+ team)","repo":"schrodinger/public_binding_free_energy_benchmark","site":"","consortium":"None","desc":"A large, openly released collection of congeneric ligand series with experimental relative binding free energies across many pharmaceutically relevant targets, plus reference FEP+ results — the standard yardstick for alchemical free energy methods.","tasks":"Relative binding free energy","note":"Reference results come from a commercial engine, but the input structures and experimental data are open and reusable.","added":"2026-07-25"},{"name":"OpenFF Protein-Ligand Benchmark","cat":"affinity","author":"Open Force Field Initiative (David Mobley, John Chodera et al.)","repo":"openforcefield/protein-ligand-benchmark","site":"","consortium":"OpenFF","desc":"Curated, ready-to-run congeneric series (structures, ligand sets, experimental affinities, prepared systems) for benchmarking free energy calculations with open force fields and open tooling end to end.","tasks":"Relative binding free energy","note":"Pairs with OpenFE for a fully open FEP stack.","added":"2026-07-25"},{"name":"SAMPL Challenges","cat":"affinity","author":"David L. Mobley, Michael K. Gilson et al.","repo":"samplchallenges/SAMPL-league","site":"https://samplchallenges.github.io","consortium":"OMSF","desc":"A long-running series of blinded prediction challenges on physical properties — hydration and partition free energies, pKa, host-guest binding, and latterly protein-ligand binding — released before the experimental answers are public.","tasks":"Solvation, logP, pKa, host-guest binding","note":"Blinded and prospective, which makes SAMPL results some of the least gameable numbers in computational chemistry.","added":"2026-07-25"},{"name":"CACHE Challenges","cat":"affinity","author":"Structural Genomics Consortium (Matthieu Schapira, Cheryl Arrowsmith et al.)","repo":"","site":"https://cache-challenge.org","consortium":"SGC","desc":"Critical Assessment of Computational Hit-finding Experiments: participants predict hits for a disease-relevant target, the organisers physically synthesise or source and experimentally test the top predictions, and results are published openly.","tasks":"Prospective hit finding","note":"The only widely-run benchmark where predictions are wet-lab tested prospectively. Slow by design, and the most honest signal in the field.","added":"2026-07-25"},{"name":"D3R Grand Challenges","cat":"affinity","author":"Drug Design Data Resource (Michael K. Gilson, Rommie Amaro et al., UC San Diego)","repo":"drugdata/D3R","site":"https://drugdesigndata.org","consortium":"D3R","desc":"A blinded challenge series (GC1-GC4 and follow-ons) on pose prediction, affinity ranking and free energy for pharmaceutical targets, with data contributed by industry before public release.","tasks":"Pose prediction, affinity ranking, FEP","note":"Now largely historical but the retrospective datasets and analyses remain a strong reference for what accuracy is realistically achievable.","added":"2026-07-25"},{"name":"GuacaMol","cat":"generative","author":"Nathan Brown, Marco Fiscato, Marwin Segler & Alain Vaucher (BenevolentAI)","repo":"BenevolentAI/guacamol","site":"","consortium":"None","desc":"The reference generative chemistry benchmark: distribution-learning tasks (validity, uniqueness, novelty, KL divergence, Fréchet ChemNet distance) plus 20 goal-directed tasks including rediscovery, similarity, isomer, median molecule and multi-property optimisation objectives.","tasks":"Distribution learning + goal-directed optimisation","note":"The goal-directed tasks are saturated — simple genetic algorithms beat most deep models. Read this alongside PMO, which adds the sample-efficiency constraint the original lacks.","added":"2026-07-25"},{"name":"MOSES (Molecular Sets)","cat":"generative","author":"Daniil Polykovskiy et al. (Insilico Medicine)","repo":"molecularsets/moses","site":"","consortium":"None","desc":"A standardised ZINC-derived training set, a reference implementation of common generative baselines (CharRNN, VAE, AAE, JT-VAE, LatentGAN) and a fixed battery of distribution-learning metrics for comparing them.","tasks":"Distribution learning","note":"Excellent for reproducibility, but distribution-learning metrics say nothing about whether generated molecules are useful. Treat as a sanity check, not an objective.","added":"2026-07-25"},{"name":"Practical Molecular Optimization (PMO)","cat":"generative","author":"Wenhao Gao, Tianfan Fu, Jimeng Sun & Connor W. Coley","repo":"wenhao-gao/mol_opt","site":"","consortium":"TDC","desc":"Re-runs 25 molecular optimisation algorithms against 23 oracles under a strict budget of 10,000 oracle calls, with hyperparameters tuned on separate tasks — reframing generative chemistry as a sample-efficiency problem.","tasks":"Sample-efficient molecular optimisation","note":"The most important methodological correction in generative chemistry benchmarking. Its headline result — that simple methods like genetic algorithms are extremely competitive — still holds up.","added":"2026-07-25"},{"name":"MolScore","cat":"generative","author":"Morgan Thomas (with Andreas Bender, Ola Engkvist et al.)","repo":"MorganCThomas/MolScore","site":"","consortium":"None","desc":"A modular scoring and benchmarking framework for de novo design that re-implements GuacaMol, MOSES and MolOpt under one interface, and lets you compose custom multi-parameter objectives with docking, QSAR models and synthesisability filters.","tasks":"Scoring, MPO, benchmark harmonisation","note":"The practical choice if you want to benchmark and deploy with the same scoring code.","added":"2026-07-25"},{"name":"Tartarus","cat":"generative","author":"AkshatKumar Nigam, Robert Pollice & Alán Aspuru-Guzik","repo":"aspuru-guzik-group/Tartarus","site":"","consortium":"None","desc":"Inverse-design benchmarks built on real physics-based simulation rather than cheap surrogate oracles — covering designing organic emitters, reaction substrates, photovoltaics and protein-ligand binders.","tasks":"Simulation-grounded inverse design","note":"Far more expensive to run than GuacaMol, which is exactly the point: the oracles cannot be trivially gamed.","added":"2026-07-25"},{"name":"DOCKSTRING","cat":"generative","author":"Miguel García-Ortegón, Gregor Simm, José Miguel Hernández-Lobato et al.","repo":"dockstring/dockstring","site":"https://dockstring.github.io","consortium":"None","desc":"A docking-based benchmark bundle: an easy Python docking wrapper, a 260k-ligand × 58-target dataset of precomputed docking scores, and three defined task families (regression, virtual screening, de novo design) with baselines.","tasks":"Docking-score regression and optimisation","note":"Makes docking-based generative benchmarks reproducible. Remember docking score is a proxy, not affinity — optimising it hard produces known pathologies.","added":"2026-07-25"},{"name":"Saturn","cat":"generative","author":"Jeff Guo & Philippe Schwaller (EPFL)","repo":"schwallergroup/saturn","site":"","consortium":"None","desc":"A sample-efficient generative design framework built on the Mamba architecture with augmented memory, benchmarked explicitly under tight oracle budgets and against synthesisability constraints.","tasks":"Sample-efficient de novo design","note":"Representative of the current generation of methods that treat oracle budget as the primary axis of comparison.","added":"2026-07-25"},{"name":"CrossDocked2020","cat":"generative","author":"Paul Francoeur & David Koes (University of Pittsburgh)","repo":"gnina/models","site":"https://bits.csb.pitt.edu/files/crossdock2020/","consortium":"None","desc":"A very large set of cross-docked protein-ligand poses across binding-site-similar targets, originally for training CNN scoring functions and now the default training and evaluation set for 3D structure-based generative models. Code and trained models live in the gnina organisation; the dataset itself is distributed separately.","tasks":"SBDD training + evaluation","note":"Widely used but widely criticised as an evaluation set for generative models — pair any CrossDocked result with PoseCheck strain and interaction analysis.","added":"2026-07-25"},{"name":"ProteinGym","cat":"structure","author":"Pascal Notin, Aaron Kollasch & Debora Marks (Harvard)","repo":"OATML-Markslab/ProteinGym","site":"https://proteingym.org","consortium":"None","desc":"The standard benchmark for protein fitness prediction: 200+ deep mutational scanning assays plus clinical variant sets, covering substitutions and indels, in zero-shot and supervised settings with a maintained leaderboard.","tasks":"Variant effect / fitness prediction","note":"Essential for biologics and protein engineering work; the zero-shot track is the cleanest comparison of protein language models available.","added":"2026-07-25"},{"name":"CASP (Critical Assessment of Structure Prediction)","cat":"structure","author":"John Moult, Krzysztof Fidelis, Andriy Kryshtafovych et al.","repo":"","site":"https://predictioncenter.org","consortium":"CASP","desc":"The biennial blinded community experiment that has defined progress in structure prediction since 1994, now including assemblies, RNA and protein-ligand complex categories relevant to drug discovery.","tasks":"Structure prediction (blinded)","note":"Not on GitHub — data and assessments are distributed through the Prediction Center. The ligand category (CASP15/16) is the part most relevant here.","added":"2026-07-25"},{"name":"CAMEO","cat":"structure","author":"Torsten Schwede lab (SIB / University of Basel)","repo":"","site":"https://cameo3d.org","consortium":"None","desc":"Continuous, fully automated weekly blinded assessment of structure prediction servers against newly released PDB entries — the rolling complement to CASP's biennial cycle, with a ligand-binding-site category.","tasks":"Continuous blinded structure assessment","note":"Because it runs weekly and automatically, it is much less susceptible to the hand-tuning that affects one-off challenge entries.","added":"2026-07-25"},{"name":"Astex Diverse Set","cat":"structure","author":"Marcel Verdonk et al. (Astex Pharmaceuticals)","repo":"","site":"","consortium":"None","desc":"85 hand-validated, pharmaceutically relevant protein-ligand complexes assembled as a clean, high-quality set for docking validation. Small, but every structure has been individually inspected.","tasks":"Docking validation","note":"Now easy for modern methods and included inside PoseBench; useful as a floor test rather than a discriminating benchmark.","added":"2026-07-25"},{"name":"ASKCOS","cat":"synthesis","author":"Connor W. Coley, Klavs Jensen & the MLPDS team (MIT)","repo":"ASKCOS/askcos-core","site":"https://askcos.mit.edu","consortium":"MLPDS","desc":"An open computer-aided synthesis planning platform covering one-step retrosynthesis, multi-step tree search, forward reaction prediction, condition recommendation and impurity prediction, with published evaluations against chemist judgement.","tasks":"Retrosynthesis, route planning, condition prediction","note":"The reference open CASP system. Route quality evaluation remains an open problem — top-k single-step accuracy correlates poorly with usable routes.","added":"2026-07-25"},{"name":"Syntheseus","cat":"synthesis","author":"Krzysztof Maziarz, Marwin Segler et al. (Microsoft Research AI4Science)","repo":"microsoft/syntheseus","site":"","consortium":"None","desc":"A benchmarking library that wraps many published single-step retrosynthesis models and search algorithms behind one interface so they can be compared fairly on multi-step route finding, exposing how much reported gains depend on evaluation details.","tasks":"Retrosynthesis model + search comparison","note":"Its accompanying analysis showed several reported rankings reverse once evaluation is standardised — read it before trusting any retrosynthesis leaderboard.","added":"2026-07-25"},{"name":"PaRoutes","cat":"synthesis","author":"Samuel Genheden & Esben Bjerrum (AstraZeneca)","repo":"MolecularAI/PaRoutes","site":"","consortium":"None","desc":"A framework for benchmarking multi-step retrosynthesis at the route level, providing reference route sets extracted from patents and metrics for comparing predicted routes against known ones.","tasks":"Multi-step route benchmarking","note":"One of the few benchmarks that evaluates whole routes rather than individual disconnections.","added":"2026-07-25"},{"name":"Open Reaction Database (ORD)","cat":"synthesis","author":"Steven M. Kearnes, Connor W. Coley et al.","repo":"open-reaction-database/ord-data","site":"https://open-reaction-database.org","consortium":"None","desc":"An open, structured schema and repository for chemical reaction data — including negative and failed reactions, which are almost entirely absent from patent-mined datasets like USPTO.","tasks":"Reaction data infrastructure","note":"Infrastructure rather than a benchmark, but it is the substrate for the next generation of honest reaction-prediction evaluation.","added":"2026-07-25"},{"name":"USPTO-50K / USPTO-MIT / USPTO-Full","cat":"synthesis","author":"Daniel Lowe (original extraction); Connor Coley, Bowen Liu et al. (splits)","repo":"connorcoley/retrosim","site":"","consortium":"None","desc":"The patent-mined reaction datasets that essentially all retrosynthesis and reaction prediction models are trained and evaluated on, in several standard sizes and splits.","tasks":"Single-step retrosynthesis, forward prediction","note":"Known issues: duplicate reactions across splits, atom-mapping artefacts that leak the answer, no failed reactions, and heavy bias toward a narrow set of reaction classes.","added":"2026-07-25"},{"name":"Open Targets Platform","cat":"target","author":"Open Targets consortium (EMBL-EBI, Wellcome Sanger, GSK, Sanofi, Genentech, Pfizer, MSD)","repo":"opentargets/platform","site":"https://platform.opentargets.org","consortium":"OpenTargets","desc":"Systematic target-disease association evidence integrated from genetics, transcriptomics, pathways, literature, drugs and animal models, with a tractability and safety layer. The standard reference for target identification and prioritisation.","tasks":"Target-disease association, tractability","note":"A resource rather than a fixed benchmark, but it is what most target-ID evaluations are scored against.","added":"2026-07-25"},{"name":"PrimeKG","cat":"target","author":"Payal Chandak, Kexin Huang & Marinka Zitnik","repo":"mims-harvard/PrimeKG","site":"","consortium":"TDC","desc":"A precision-medicine knowledge graph integrating 20 resources into ~4M relationships across 17,000+ diseases, spanning genes, drugs, phenotypes, pathways and exposures — built for drug repurposing and target reasoning benchmarks.","tasks":"Knowledge graph reasoning, repurposing","note":"Now a standard substrate for evaluating LLM and GNN reasoning over biomedical knowledge.","added":"2026-07-25"},{"name":"GEARS","cat":"target","author":"Yusuf Roohani, Kexin Huang & Jure Leskovec (Stanford)","repo":"snap-stanford/GEARS","site":"","consortium":"None","desc":"A benchmark and method for predicting transcriptional outcomes of single and combinatorial genetic perturbations, including generalisation to perturbations never seen during training.","tasks":"Perturbation response prediction","note":"Subsequent work has questioned whether models beat simple mean-expression baselines on some splits — check the baseline before believing an improvement.","added":"2026-07-25"},{"name":"Open Problems in Single-Cell Analysis","cat":"target","author":"Open Problems community (Daniel Burkhardt, Robrecht Cannoodt et al.)","repo":"openproblems-bio/openproblems","site":"https://openproblems.bio","consortium":"OpenProblems","desc":"Formalised, continuously re-run benchmarks for single-cell tasks (batch integration, denoising, label projection, perturbation prediction, spatial decomposition) with living leaderboards rather than frozen paper results.","tasks":"Single-cell method benchmarking","note":"The continuous re-running model is what most drug discovery benchmarks should aspire to.","added":"2026-07-25"},{"name":"scPerturb","cat":"target","author":"Stefan Peidli, Nils Blüthgen et al.","repo":"sanderlab/scPerturb","site":"http://projects.sanderlab.org/scperturb/","consortium":"None","desc":"Harmonised single-cell perturbation datasets (genetic and chemical) with uniform processing and distance-based evaluation metrics, enabling like-for-like comparison across labs and platforms.","tasks":"Perturbation response, E-distance evaluation","note":"The E-distance metric it introduced is a more honest measure than per-gene correlation.","added":"2026-07-25"},{"name":"OpenBioLink","cat":"target","author":"Anna Breit, Matthias Samwald et al.","repo":"OpenBioLink/OpenBioLink","site":"","consortium":"None","desc":"A biomedical knowledge graph benchmark for link prediction with explicitly leakage-controlled train/test splits and both directed and undirected variants — built because earlier biomedical KG benchmarks were trivially leaky.","tasks":"Biomedical link prediction","note":"Its main contribution is the negative sampling and split methodology.","added":"2026-07-25"},{"name":"Clinical Trial Outcome Prediction (HINT / TOP)","cat":"clinical","author":"Tianfan Fu, Kexin Huang, Cao Xiao & Jimeng Sun","repo":"futianfan/clinical-trial-outcome-prediction","site":"","consortium":"TDC","desc":"The Trial Outcome Prediction benchmark: ~17,000 trials with drug molecules, disease codes and eligibility criteria, split by phase, for predicting whether a trial will succeed. Distributed through TDC.","tasks":"Phase I/II/III outcome prediction","note":"The hardest kind of task to benchmark honestly — outcome labels are noisy, and the strongest signal often comes from sponsor and indication metadata rather than the molecule.","added":"2026-07-25"},{"name":"TrialBench","cat":"clinical","author":"Jintai Chen, Tianfan Fu et al.","repo":"ML2Health/ML2ClinicalTrials","site":"","consortium":"None","desc":"A multi-modal clinical trial benchmark suite spanning trial duration, patient dropout, serious adverse event and mortality prediction as well as outcome, with standardised splits across phases.","tasks":"Multi-task clinical trial prediction","note":"Broader than TOP; the adverse-event tasks are the most clinically actionable.","added":"2026-07-25"},{"name":"MIMIC-IV","cat":"clinical","author":"Alistair Johnson, Leo Anthony Celi, Roger Mark et al. (MIT LCP)","repo":"MIT-LCP/mimic-code","site":"https://mimic.mit.edu","consortium":"None","desc":"De-identified ICU and hospital EHR data with a large open codebase of derived concepts and benchmark task definitions. The reference dataset for clinical outcome modelling and real-world evidence method development.","tasks":"Clinical outcome prediction, RWE","note":"Requires credentialed access and training. Single-centre, so external validity is limited — an important caveat for drug-effect estimation.","added":"2026-07-25"},{"name":"LAB-Bench","cat":"agent","author":"Jon M. Laurent, Joseph D. Janizek, Andrew D. White et al. (FutureHouse)","repo":"Future-House/LAB-Bench","site":"","consortium":"FutureHouse","desc":"Over 2,400 multiple-choice questions covering practical biology research tasks: literature search and reasoning, figure and table interpretation, protocol planning, DNA and protein sequence manipulation, and database access.","tasks":"Biology research assistance","note":"Compares model performance against human expert baselines, which makes the results interpretable rather than just relative.","added":"2026-07-25"},{"name":"BixBench","cat":"agent","author":"Ludovico Mitchener et al. (FutureHouse)","repo":"Future-House/BixBench","site":"","consortium":"FutureHouse","desc":"An agentic bioinformatics benchmark built from real analysis notebooks: agents must explore real datasets, write and execute code, and answer open-ended analytical questions rather than pick from options.","tasks":"Agentic bioinformatics analysis","note":"Open-answer format makes it much harder than multiple choice; frontier model scores were strikingly low at release.","added":"2026-07-25"},{"name":"ChemBench","cat":"agent","author":"Adrian Mirza, Nawaf Alampara & Kevin Maik Jablonka (LamaLab)","repo":"lamalab-org/chembench","site":"https://chembench.org","consortium":"None","desc":"A large chemistry evaluation corpus (2,700+ curated question-answer pairs) spanning general, organic, inorganic, analytical, physical and toxicological chemistry, benchmarked against human chemists.","tasks":"Chemistry knowledge and reasoning","note":"Its finding that models exceed average human chemists overall while failing on specific safety-relevant reasoning is the interesting part, not the headline score.","added":"2026-07-25"},{"name":"BioDiscoveryAgent","cat":"agent","author":"Yusuf Roohani, Andrew Lee, Jure Leskovec et al. (Stanford)","repo":"snap-stanford/BioDiscoveryAgent","site":"","consortium":"None","desc":"An agent and evaluation setting for designing genetic perturbation experiments — the agent iteratively proposes which genes to perturb next, evaluated on how efficiently it discovers hits in real Perturb-seq screens.","tasks":"Closed-loop experimental design","note":"One of the few benchmarks that measures discovery efficiency in a loop rather than one-shot prediction accuracy.","added":"2026-07-25"},{"name":"PaperQA / LitQA","cat":"agent","author":"Michael Skarlinski, Andrew D. White et al. (FutureHouse)","repo":"Future-House/paper-qa","site":"","consortium":"FutureHouse","desc":"A retrieval-augmented agent for scientific literature plus the LitQA evaluation set of questions answerable only from the full text of recent papers, designed to test genuine literature retrieval rather than memorisation.","tasks":"Literature-grounded question answering","note":"Because questions target recent full texts, contamination is much less of a concern than in general science QA sets.","added":"2026-07-25"},{"name":"ChemCoTBench","cat":"agent","author":"IDEA Research (IDEA-XL)","repo":"IDEA-XL/ChemCoTBench","site":"","consortium":"None","desc":"A chemistry reasoning benchmark with chain-of-thought supervision covering molecule editing, optimisation and reaction prediction, aimed at evaluating multi-step chemical reasoning rather than recall.","tasks":"Chemical reasoning, molecule editing","note":"Newer and less battle-tested than ChemBench; useful as a complement.","added":"2026-07-25"},{"name":"TDC Drug-Target Interaction Benchmarks","cat":"affinity","author":"Kexin Huang & Marinka Zitnik","repo":"mims-harvard/TDC","site":"https://tdcommons.ai/benchmark/dti_dg_group/overview/","consortium":"TDC","desc":"Drug-target binding affinity tasks (BindingDB, DAVIS, KIBA) plus a domain-generalisation group that splits by patent year to test whether models extrapolate to genuinely new chemistry and targets.","tasks":"Drug-target affinity, domain generalisation","note":"The DTI-DG time-split group is far more informative than the random-split DAVIS/KIBA numbers most papers report.","added":"2026-07-25"},{"name":"DeepChem Benchmark Suite","cat":"platform","author":"Bharath Ramsundar & the DeepChem community","repo":"deepchem/deepchem","site":"https://deepchem.io","consortium":"DeepChem","desc":"The library that hosts MoleculeNet loaders alongside a wide set of featurisers, splitters (random, scaffold, butina, stratified) and reference model implementations — the practical way most people actually run these benchmarks.","tasks":"Featurisation, splitting, reference models","note":"Its splitter implementations are the de facto standard; use them rather than rolling your own scaffold split.","added":"2026-07-25"},{"name":"Open Force Field Benchmarks","cat":"affinity","author":"Open Force Field Initiative","repo":"openforcefield/openff-benchmark","site":"https://openforcefield.org","consortium":"OpenFF","desc":"Systematic benchmarking of small-molecule force fields against quantum chemistry reference data — conformer energies, geometries and torsion profiles — which underpins the reliability of every downstream physics-based affinity prediction.","tasks":"Force field validation","note":"Easy to overlook, but force field error is often the dominant error term in FEP results.","added":"2026-07-25"},{"name":"MUV (Maximum Unbiased Validation)","cat":"affinity","author":"Sebastian G. Rohrer & Knut Baumann","repo":"","site":"","consortium":"None","desc":"PubChem-derived virtual screening datasets constructed with a spatial statistics criterion so that actives are maximally spread among decoys, explicitly removing analogue bias and artificial enrichment.","tasks":"Unbiased virtual screening","note":"Extremely hard by design and included in MoleculeNet. Near-random performance is normal, which makes small reported gains suspicious.","added":"2026-07-25"},{"name":"CardioTox / hERG Benchmarks","cat":"admet","author":"Abdul Karim et al.","repo":"Abdulk084/CardioTox","site":"","consortium":"None","desc":"A hERG cardiotoxicity benchmark with explicitly defined external test sets chosen to be structurally distinct from training data, addressing the near-duplicate problem in earlier hERG datasets.","tasks":"hERG blockade / cardiotoxicity","note":"hERG is one of the few endpoints where public data is large enough to be genuinely useful, but assay heterogeneity across sources is severe.","added":"2026-07-25"},{"name":"TDC Molecule Generation Oracles","cat":"generative","author":"Wenhao Gao, Tianfan Fu & Kexin Huang","repo":"mims-harvard/TDC","site":"https://tdcommons.ai/functions/oracles/","consortium":"TDC","desc":"17+ standardised optimisation oracles (GSK3β, JNK3, DRD2, QED, SA, docking scores, MPO objectives) with fixed implementations, so that reported optimisation results are comparable across papers.","tasks":"Generative optimisation objectives","note":"The fixed implementations solved a real reproducibility problem — small differences in oracle definitions used to make results incomparable.","added":"2026-07-25"},{"name":"TDC Single-Cell & Perturbation Benchmarks (TDC-2)","cat":"target","author":"Marinka Zitnik and TDC contributors","repo":"mims-harvard/TDC","site":"https://tdcommons.ai","consortium":"TDC","desc":"The TDC-2 extension covering contextual single-cell tasks, counterfactual perturbation prediction, and multimodal model-hub evaluation, bringing target-side biology into the same framework as the chemistry tasks.","tasks":"Single-cell, perturbation, multimodal","note":"The contextual single-cell and multimodal work described here has since moved into PyTDC, which forked from this repository and is developed separately. For current single-cell tasks and model-server tooling, see the PyTDC entry; this entry covers what remains in the mims-harvard/TDC dataset store.","added":"2026-07-25"},{"name":"DDD Benchmark as a Service (Insilico)","cat":"platform","author":"Insilico Medicine (Alex Zhavoronkov et al.)","repo":"","site":"https://dddbench.insilico.com","consortium":"Insilico","desc":"Launched 30 July 2026 and billed as the industry's first drug discovery and development benchmark offered as a service. Any model exposed through a standard chat-completions API can be submitted; Insilico scores the outputs against expert reference baselines drawn from its own validated programmes and returns a standardised scorecard against leading models. The headline task, Drug Candidate Essentials, runs an end-to-end programme from hit identification through preclinical candidate nomination.","tasks":"End-to-end discovery, hit ID to PCC nomination, model evaluation","note":"Built explicitly to attack data contamination — the argument being that public benchmarks inflate scores because test items already sit in training corpora. The trade-off is that the reference baselines are proprietary, so you are trusting Insilico's curation rather than inspecting it. That is the same bargain as any blinded challenge, and arguably a fair one, but worth stating plainly.","added":"2026-07-30"},{"name":"Drug Discovery Benchmark (DDB) Leaderboard","cat":"platform","author":"Insilico Medicine","repo":"","site":"https://ddb.insilico.com/","consortium":"Insilico","desc":"The public leaderboard portal for end-to-end drug discovery tasks, covering LLM performance on clinical trial prediction and target discovery among others. One of three portals released alongside the MMAI Gym expansion.","tasks":"Clinical trial prediction, target discovery, LLM evaluation","note":"Public-facing scores; the underlying held-out items are not distributed, which is the point but also limits independent reproduction.","added":"2026-07-30"},{"name":"Science MMAI Gym","cat":"platform","author":"Insilico Medicine","repo":"","site":"https://insilico.com/mmai","consortium":"Insilico","desc":"A post-training and benchmarking environment for scientific AI, expanded in 2026 with three public leaderboard portals: ScienceAI Bench (broad scientific reasoning across biology, chemistry, longevity, materials science and agriculture), Drug Discovery Benchmark, and Insilico Bench (proprietary benchmarks for drug discovery and other complex scientific problems).","tasks":"Scientific reasoning, drug discovery, post-training environments","note":"Positioned as a training environment as much as an evaluation suite — useful to keep the two roles distinct when reading results, since a gym you train in is not a clean test set.","added":"2026-07-30"},{"name":"TargetBench (TargetPro–TargetBench)","cat":"target","author":"Insilico Medicine","repo":"","site":"https://insilico.com/news/77g216v1e1-insilico-medicine-advances-ai-driven-tar","consortium":"Insilico","desc":"A benchmark for target identification capability, released as the evaluation half of a paired framework with the TargetPro target discovery model. Pairs prediction with benchmarking so that a target-ID claim is scored rather than asserted.","tasks":"Target identification, target prioritisation","note":"Target ID is notoriously hard to benchmark because ground truth is slow and contested. Read the choice of positive set carefully — it largely determines the result.","added":"2026-07-30"},{"name":"ChemCensor / CREED / URSA-expert-2026","cat":"synthesis","author":"Insilico Medicine, Generative AI & Quantum Computing R&D Center (Abu Dhabi)","repo":"","site":"https://arxiv.org/abs/2602.03554","consortium":"Insilico","desc":"An ICML 2026 framework arguing that single-answer Top-K accuracy on USPTO-50K is fundamentally unsuited to retrosynthesis, because real chemistry is multi-solution. Introduces ChemCensor, a chemistry-aware metric scoring plausibility via reaction centres and functional groups; CREED, a 6.4M-reaction ChemCensor-validated dataset; and URSA-expert-2026, an expert-annotated out-of-domain set of 100 novel targets built to be leakage-free.","tasks":"Single-step retrosynthesis, chemical plausibility","note":"The critique of ground-truth Top-K scoring is the substantive contribution and applies well beyond this framework. Supporting material was promised on Zenodo, Hugging Face and GitHub — check for a repository before relying on it.","added":"2026-07-30"},{"name":"SupraBench","cat":"admet","author":"See arXiv 2606.13477","repo":"","site":"https://arxiv.org/pdf/2606.13477","consortium":"None","desc":"A 2026 benchmark for supramolecular chemistry, covering host-guest and non-covalent association tasks that fall outside the covalent, single-molecule framing of MoleculeNet and TDC property prediction.","tasks":"Supramolecular chemistry, host-guest binding","note":"Fills a genuine gap, but new and not yet widely adopted — treat baselines as provisional.","added":"2026-07-30"},{"name":"SMDD-Bench","cat":"agent","author":"See arXiv 2605.21740","repo":"","site":"https://arxiv.org/pdf/2605.21740","consortium":"None","desc":"Evaluates LLM agents on real-world small molecule drug design tasks, moving past single-turn question answering toward multi-step design problems that resemble what a medicinal chemist actually faces.","tasks":"Agentic small molecule design","note":"Part of a fast-growing cluster of agentic design benchmarks; cross-check against PMO-style oracle-budget evaluation before concluding an agent is efficient.","added":"2026-07-30"},{"name":"LAB-Bench 2","cat":"agent","author":"FutureHouse","repo":"Future-House/LAB-Bench","site":"https://arxiv.org/html/2604.09554v2","consortium":"FutureHouse","desc":"The 2026 successor to LAB-Bench, revised to address saturation and known weaknesses in the original multiple-choice format while retaining human-expert baselines for biology research tasks.","tasks":"Biology research assistance, literature reasoning","note":"Prefer this over the original where both are available — several LAB-Bench v1 categories were approaching ceiling.","added":"2026-07-30"},{"name":"BioAgent Bench","cat":"agent","author":"See arXiv 2601.21800","repo":"","site":"https://arxiv.org/pdf/2601.21800","consortium":"None","desc":"An agent evaluation suite for bioinformatics workflows, requiring agents to execute real analyses end to end rather than answer questions about them.","tasks":"Agentic bioinformatics","note":"Overlaps substantially with BixBench; compare task construction before treating the two as independent evidence.","added":"2026-07-30"},{"name":"PoseX","cat":"docking","author":"CataAI","repo":"CataAI/PoseX","site":"https://arxiv.org/abs/2505.01700","consortium":"None","desc":"A docking benchmark built around cross-docking rather than the easier self-docking case, with 718 self-docking and 1,312 cross-docking entries and a real-time leaderboard. It evaluates 23 methods side by side across three families — physics-based (Glide), AI docking (DiffDock) and AI co-folding (AlphaFold3, Boltz, Chai) — which most docking benchmarks do not do in one place.","tasks":"Self-docking, cross-docking, pose prediction","note":"The headline claim that AI beats physics is partly an artefact of how the cross-docking pairs were selected — check the holdout construction against PoseBusters-style physical-validity filters before repeating it. The authors themselves flag ligand chirality failures in most co-folding models, with Boltz-1x the exception.","added":"2026-07-31"},{"name":"CoFD-Bench","cat":"docking","author":"Zhang et al., Acta Pharmacologica Sinica","repo":"","site":"https://www.nature.com/articles/s41401-025-01721-5","consortium":"None","desc":"A covalent-complex benchmark of 218 recently resolved structures, comparing classical covalent docking (AutoDock-GPU, CovDock, GNINA) against co-folding models (AlphaFold3, Chai-1, Boltz-1x). Covalent binders are systematically excluded from most docking benchmarks, so this fills a real gap.","tasks":"Covalent pose prediction, interaction recovery","note":"Small by modern standards at 218 complexes, and 'recently resolved' is not the same as a clean time split — some of these structures may still predate the training cutoffs of the co-folding models being tested. No public repository located; evaluation code appears to live only in the supplementary material.","added":"2026-07-31"},{"name":"AssayBench","cat":"target","author":"De Brouwer et al., Genentech","repo":"Genentech/AssayBench","site":"https://arxiv.org/abs/2605.10876","consortium":"None","desc":"Frames virtual-cell prediction as gene-rank recovery across 1,920 public CRISPR screens spanning five phenotype classes, scored with an adjusted nDCG that is comparable across assays of very different sizes. Unlike most perturbation benchmarks it evaluates LLMs and agents directly from a textual screen description rather than requiring an expression-matrix model.","tasks":"Phenotypic screen prediction, gene ranking","note":"Built entirely from published BioGRID screens, so memorisation is a live concern — the authors include a memorisation analysis and a year-based split precisely because performance correlates with publication year and citation count. Use the year split, not the random one.","added":"2026-07-31"},{"name":"VCBench","cat":"target","author":"See bioRxiv 2026.06.18.733146","repo":"","site":"https://www.biorxiv.org/content/10.64898/2026.06.18.733146v1","consortium":"None","desc":"Consolidates four separate virtual-cell evaluation frameworks into seven capability dimensions — perturbation response, cross-species transfer, GRN inference, modality integration, temporal dynamics, multi-scale integration and in-silico experimentation. Its distinguishing feature is pre-registered linear and nearest-neighbour baselines, so a foundation model has to beat a trivial control to count as working.","tasks":"Single-cell foundation model evaluation","note":"Only five dimensions of the seven are actually testable today, and the headline result is that the evaluated foundation models (Geneformer, scGPT, UCE, TranscriptFormer, Arc State) frequently fail to beat the simple baselines. Preprint, not yet independently reproduced, and no public repository confirmed.","added":"2026-07-31"},{"name":"BiomniBench","cat":"agent","author":"See bioRxiv 2026.05.12.724604","repo":"","site":"https://huggingface.co/datasets/phylobio/BiomniBench-DA","consortium":"None","desc":"Scores the whole agent trajectory against expert-authored, task-specific rubrics instead of only the final answer, making it the first process-level benchmark for biomedical research agents. The BiomniBench-DA instantiation has 100 multi-step data-analysis tasks derived from Nature/Cell/Science papers, each co-developed with an original author or domain expert.","tasks":"Agentic data analysis, process-level scoring","note":"Only 50 of the 100 tasks are public; the rest are held out as a contamination-resistant set, so self-reported numbers on the public half are not comparable to the official ones. Rubric-based grading is LLM-judged, which introduces its own scorer bias.","added":"2026-07-31"},{"name":"BioML-bench","cat":"agent","author":"Science Machine","repo":"science-machine/biomlbench","site":"https://www.biorxiv.org/content/10.1101/2025.09.01.673319v2","consortium":"None","desc":"An MLE-bench-derived harness that makes agents build complete ML solutions end to end across protein engineering, drug discovery, single-cell omics, medical imaging and clinical biomarkers, with human baselines for each task. Agents run containerised with RDKit and BioPython preinstalled, so it measures ML engineering ability rather than recall.","tasks":"End-to-end biomedical ML, agentic modelling","note":"Self-described v0.1-alpha and explicitly warns of bugs and incomplete features. Several tasks are drawn from Polaris and Kaggle competitions whose solutions are public, so leaderboard leakage is plausible for any agent with web access — run it sandboxed.","added":"2026-07-31"},{"name":"OpenADMET Blind Challenges (PXR)","cat":"admet","author":"OpenADMET","repo":"","site":"https://openadmet.org/blindchallenges/","consortium":"OMSF","desc":"A rolling series of prospective blind challenges on ADMET endpoints, where the test data is genuinely unseen because it has not been measured yet at submission time. The current round targets human PXR induction with both an activity-prediction and a structure-prediction track; the previous ExpansionRx round covered nine endpoints from a real RNA-targeting campaign and closed in January 2026.","tasks":"Prospective ADMET prediction, PXR induction","note":"Prospective design removes the leakage problem that plagues retrospective ADMET benchmarks, but each round is a single chemical series from a single campaign, so ranking well says little about generalisation. Leaderboards are round-scoped and hosted on Hugging Face rather than archived in one place, which makes historical comparison awkward.","added":"2026-07-31"},{"name":"TxBench-PP","cat":"agent","author":"Workman et al., Latch Bio","repo":"latchbio/txbench-pp","site":"https://arxiv.org/abs/2606.19245","consortium":"None","desc":"One hundred verifiable evaluations built from real preclinical assay data, indexed by program stage and assay type, covering mechanism-of-action and PD reasoning, target engagement, causal target validation, developability and translational efficacy. The design deliberately tests whether an agent can reach the right conclusion from the data rather than recall the published answer. Across 16 model-harness configurations and 4,800 trajectories no system passed reliably; the best reached roughly 59% of endpoint attempts.","tasks":"MoA and PD reasoning, target engagement, developability, translational efficacy","note":"Only seven of the 100 evaluations are public in the repo — the rest are held out, so third parties cannot reproduce the headline numbers or audit the grading rubric. Sourced from a single company's internal program data, and pass/fail scoring collapses partially correct pharmacological reasoning into a binary.","added":"2026-08-01"},{"name":"ToxReason","cat":"admet","author":"DMIS Lab, Korea University (ACL 2026 Findings)","repo":"","site":"https://arxiv.org/abs/2604.06264","consortium":"None","desc":"Grades not just whether a model predicts liver, heart or kidney toxicity but whether the stated mechanism is right, by grounding each instance in an Adverse Outcome Pathway from molecular initiating event through to adverse outcome. Built by joining curated AOP knowledge with experimental drug-target interaction evidence and chemical-toxicity labels for 193 chemicals. Almost every other tox benchmark scores the label alone, which lets a model be right for biologically fluent nonsense reasons.","tasks":"Organ toxicity prediction, mechanistic pathway reasoning","note":"193 chemicals is small, and coverage is bounded by what the AOP knowledge bases already contain — it rewards recall of documented pathways and cannot detect novel mechanisms. The paper itself finds predictive accuracy and reasoning quality often diverge, so a single headline score is misleading. No repository confirmed at time of indexing; the paper states code is released but the location could not be verified.","added":"2026-08-01"},{"name":"AbBiBench","cat":"structure","author":"MSBMI-SAFE and collaborators","repo":"MSBMI-SAFE/AbBiBench","site":"https://arxiv.org/abs/2506.04235","consortium":"None","desc":"Around 186,000 experimental affinity measurements of antibody mutants across 13 antibodies and 9 antigens including influenza, HER2, VEGF, integrin, Ang2 and SARS-CoV-2, with structures for the complexes. Its key move is scoring the whole antibody-antigen complex rather than the antibody in isolation, which is how most protein language model evaluations were previously set up. Fifteen protein models are compared spanning masked and autoregressive LMs, inverse folding, diffusion and geometric models.","tasks":"Affinity maturation ranking, binding-aware antibody design","note":"Nine antigens is a narrow slice of antigen space and several are heavily studied targets likely present in model training data. Deep mutational scanning measurements are assay-specific and not directly comparable across the source studies, so cross-antibody aggregate scores blur real differences in dynamic range.","added":"2026-08-01"},{"name":"CHIMERA-Bench","cat":"structure","author":"Mansoor et al. (GEM workshop, ICLR 2026)","repo":"","site":"https://arxiv.org/abs/2603.13431","consortium":"None","desc":"A single canonical task — epitope-conditioned CDR sequence-structure co-design — over 2,922 deduplicated antibody-antigen complexes with epitope and paratope annotations. Three splits test generalisation to unseen epitopes, unseen antigen folds and prospective temporal targets, and the protocol adds epitope-specificity metrics on top of the usual AAR, RMSD and DockQ. Antibody design papers have historically each used their own test set, which is what makes the eleven-method head-to-head here worth having.","tasks":"Epitope-conditioned CDR co-design, antigen-fold generalisation","note":"Workshop paper, not yet peer reviewed at full-conference length, and no repository could be confirmed at time of indexing. All metrics are computational proxies — amino acid recovery and DockQ against a single crystal structure penalise valid alternative binders, and nothing here is validated by experiment.","added":"2026-08-01"},{"name":"BioDesignBench","cat":"agent","author":"Kim and Romero (Romero Lab)","repo":"","site":"https://www.biorxiv.org/content/10.64898/2026.05.06.723381v1","consortium":"None","desc":"Seventy-six expert-curated protein design tasks crossed over molecular subject (antibody, enzyme, mini-protein binder, scaffold, fluorescent protein) and design intent (de novo versus redesign), with a public leaderboard. Unusually it ships human and non-LLM baselines alongside the agent scores, plus behavioural metrics derived from tool-use traces rather than final answers only. That trace-level analysis is what separates it from the growing pile of outcome-only agent benchmarks.","tasks":"De novo protein design, redesign, tool-use behaviour analysis","note":"Scores are computational judgements of designs that were never expressed or assayed, so a high score means the agent produced something a scoring function likes. Seventy-six tasks is small enough that a handful of items can move the ranking, and the expert-curated rubric has not been independently re-scored. Repository location not confirmed at time of indexing.","added":"2026-08-01"},{"name":"Novelty-Tiered Affinity Benchmark (NTAB)","cat":"affinity","author":"Mattsson (Enlace Bio) and Walters (OpenADMET)","repo":"bamattsson/ntab","site":"https://www.biorxiv.org/content/10.64898/2026.06.29.735309v1","consortium":"None","desc":"A binding-affinity benchmark that partitions its test set into ligand-novelty tiers by Tanimoto similarity rather than splitting on protein sequence identity. The accompanying analysis shows why the usual approach fails: 'target mirroring' means homologous proteins with sequence identity as low as 0.2 still have correlated binding profiles, and a ligand-only baseline with no structural information reaches r = 0.66 on FEP+ 4 and r = 0.36 on OpenFE. In the hardest tier (similarity < 0.35) that same baseline drops to r = 0.14, which is the point — it gives a floor against which co-folding models like Boltz-2 can be judged honestly.","tasks":"Binding affinity prediction, generalisation to novel chemotypes, leakage-controlled evaluation","note":"This is a critique paper first and a benchmark second, so the evaluation set is smaller than the FEP+ and OpenFE collections it is arguing against and the tiers inherit whatever assay heterogeneity ChEMBL 36 carries. Novelty is defined by 2D fingerprint similarity, which is a crude proxy — a scaffold hop with low Tanimoto can still be biologically familiar to a model. Preprint, not yet independently reproduced.","added":"2026-08-04"},{"name":"MolRGen","cat":"generative","author":"Formont, Darrin, Ben Ayed and Piantanida (Paris-Saclay, ETS Montreal, Mila, Mistral AI)","repo":"","site":"https://arxiv.org/abs/2603.18256","consortium":"None","desc":"Roughly 4,500 protein-pocket targets expanded into about 50k multi-objective prompts that combine docking score with QED, synthetic accessibility, logP and physicochemical descriptors, shipped with an open-source verifier that computes the reward at generation time. That last part is the distinguishing move: unlike caption-based or molecule-editing benchmarks it needs no reference molecule, so the same artifact serves as both an evaluation harness and an RL training environment. It also adds a diversity-aware top-k metric that penalises models for finding one good scaffold and stopping.","tasks":"De novo multi-objective molecular generation, property prediction, RL reward verification","note":"The reward is a docking score, so the benchmark inherits every known failure of docking as an affinity proxy — a model can top the leaderboard by learning to exploit the scoring function. Models see only textual descriptions of targets, not 3D pocket geometry, which caps how much of the task is really structure-based. The authors' own GRPO fine-tuning result shows scores improving while chemical diversity collapses, which is worth reading as a warning about the metric. No repository path confirmed at time of indexing.","added":"2026-08-04"},{"name":"PromptBio-Bench","cat":"agent","author":"PromptBio Inc.","repo":"PromptBio/promptbio-bench","site":"https://www.biorxiv.org/content/10.64898/2026.05.05.723092v2","consortium":"None","desc":"244 expert-curated bioinformatics and data-science tasks graded at three difficulty levels, scored by structured file comparison against expert reference answer files rather than by free-text judging. The file-level grading is what separates it from the QA-style computational-biology benchmarks — an agent has to produce an artifact that matches, not an answer that reads well. Reported alongside accuracy are wall-clock time and token consumption per run, so cost is visible next to capability.","tasks":"End-to-end bioinformatics analysis, multi-step data science workflows, agent cost accounting","note":"Every author is affiliated with PromptBio Inc., and ToolsGenie — one of only three agents evaluated — is a PromptBio product, so the headline comparison is a vendor scoring its own tool. Similarity-to-reference scoring punishes valid alternative analyses that reach the same conclusion by a different route. Three agents is a thin comparison set and the tasks have not been independently re-curated.","added":"2026-08-04"},{"name":"CT Open","cat":"clinical","author":"Cao et al.","repo":"","site":"https://arxiv.org/abs/2604.16742","consortium":"None","desc":"A live clinical-trial outcome prediction challenge running four quarterly windows a year, where predictions must be submitted before the window opens and are scored only on trials that had no public results beforehand but acquired them during it. That forward-only design makes it one of the few benchmarks that can fairly evaluate an agent with unrestricted web search, since there is nothing to retrieve. Questions are posed at outcome-measure level against specific study arms rather than as a single trial-level success or failure label, which is finer-grained than HINT/TOP or TrialBench.","tasks":"Prospective clinical trial outcome prediction, outcome-measure level forecasting, web-enabled agent evaluation","note":"Decontamination depends on an automated LLM web-search pipeline that dates the earliest public mention of an outcome — if that pipeline misses a mention, a contaminated question enters the scored set, and its error rate is not externally audited. Each window is limited to whatever trials happen to report in that quarter, so sample sizes are modest and quarter-to-quarter results are not directly comparable. Being prospective, it also cannot be re-run offline to reproduce a past leaderboard.","added":"2026-08-04"},{"name":"PepBenchmark","cat":"admet","author":"ZGCI AI4S Peptide group","repo":"ZGCI-AI4S-Pep/PepBenchmark","site":"https://arxiv.org/abs/2604.10531","consortium":"None","desc":"Peptides have been the conspicuous gap in this index: every major property-prediction benchmark here is built for small molecules, and peptide papers have each rolled their own splits. PepBenchmark ships 29 canonical and 6 non-canonical peptide datasets across 7 task groups, a fixed preprocessing and splitting pipeline, and a leaderboard with baselines from four method families — fingerprint, GNN, protein language model and SMILES-based — so the representation question can be answered on common ground rather than per paper.","tasks":"Peptide property and activity prediction, representation comparison","note":"Peptide datasets are small and heavily skewed toward antimicrobial and cell-penetrating activity, so aggregate scores are dominated by a few well-populated tasks. Non-canonical peptides are only six datasets, which is exactly the regime where the SMILES and fingerprint baselines should matter most and where the evidence is thinnest. New and not yet independently reproduced.","added":"2026-08-07"},{"name":"Tox21 Reproducible Leaderboard","cat":"admet","author":"Ebner et al. (ML-JKU, Johannes Kepler University Linz)","repo":"","site":"https://huggingface.co/spaces/ml-jku/tox21_leaderboard","consortium":"None","desc":"Restores the original Tox21 Challenge data and evaluation as a live, automated leaderboard on Hugging Face, with the untouched challenge training set published as a companion dataset. The motivating observation is specific and damning: as Tox21 was absorbed into MoleculeNet, TDC and OGB the labels were altered and imputed, so a decade of reported Tox21 numbers are not comparable to each other or to the 2014 challenge results. Inference and leaderboard updates run automatically through HF Datasets and Spaces.","tasks":"12 Tox21 assay endpoints, reproducible toxicity evaluation","note":"Fixes comparability but not the underlying data — Tox21 is still a 2014 assay panel with heavy class imbalance and the label noise that comes with high-throughput screening. Automated HF submission also means results are self-reported by whoever submits, with no independent re-run of the training pipeline behind a score.","added":"2026-08-07"},{"name":"DrugPlayGround","cat":"agent","author":"See bioRxiv 2026.04.04.716470 / arXiv 2604.02346","repo":"","site":"https://www.biorxiv.org/content/10.64898/2026.04.04.716470v1","consortium":"None","desc":"Splits LLM evaluation for drug discovery into two branches that are usually conflated: the quality of generated text describing physicochemical properties, drug synergy, drug-protein interactions and perturbation response, and separately the downstream usefulness of the model's embeddings on the same tasks. Scoring combines automatic metrics with chemist review, which is rare in this category.","tasks":"LLM text generation and embedding quality for drug tasks","note":"Nearly all source facts are drawn from public resources that sit in pretraining corpora, so this measures recall and articulation more than reasoning. Chemist-in-the-loop scoring is the strongest part and also the least reproducible — the panel size and inter-rater agreement determine how much weight those numbers carry. Preprint, no public repository confirmed at time of indexing.","added":"2026-08-07"},{"name":"ChemCost","cat":"agent","author":"Wu, Huang, Shen et al. (Carnegie Mellon / Notre Dame; Isayev and Zhang groups)","repo":"","site":"https://arxiv.org/abs/2605.07251","consortium":"None","desc":"Asks an agent to price a reaction: ground ambiguous chemical names, retrieve supplier quotes, pick valid purchasable packs under purity and quantity rules, normalise stoichiometry and aggregate cost per gram of product. 1,427 evaluable reactions drawn from ORD, named-reaction textbooks, Organic Syntheses, PaRoutes and ChemPU, with ground truth computed deterministically from a frozen snapshot of 230,775 supplier quotes over 2,261 chemicals. The point of the design is that it is judge-free — every intermediate stage (grounding, retrieval, pack selection, arithmetic) is checked automatically, so failures can be attributed rather than guessed at. A four-layer noise-injection suite perturbs aliases, quantities, missing fields and formatting while preserving the answer. Cost is the axis that route-planning benchmarks in this index systematically ignore.","tasks":"Chemical procurement cost estimation, tool-use grounding, stage-level agent diagnosis","note":"The strongest agents reach only 50.6% accuracy within 25% relative error on clean inputs, which is the headline finding and also a warning that the metric band is wide. The task is deliberately abstracted from real procurement: fixed 1 mmol scale, 95% purity threshold, one pack-selection rule, a frozen price snapshot, and no modelling of stock, shipping, lead time, regional pricing or bulk discounts — so it measures structured tool-use discipline more than commercial judgement. Experiments cover only ReAct-style agents with a fixed tool set and step budget. No public repository confirmed at time of indexing; the supplier database is the artifact everything depends on and its availability was not verifiable.","added":"2026-08-10"},{"name":"HEDGEHOG","cat":"generative","author":"Ryabchenko, Gurevich, Pak et al. (Ligand Pro; Skoltech AI Center)","repo":"","site":"https://arxiv.org/abs/2607.13155","consortium":"None","desc":"Evaluates molecular generators as an industrial hit-identification funnel rather than a scoreboard of distribution metrics: six sequential stages run preprocessing and standardisation, physicochemical descriptor screening, structural alerts and graph-sanity checks, synthesis feasibility, docking and binding affinity, then 3D pose and interaction checks. Cheap filters run first so expensive docking is only spent on molecules that survive, and the survival curve across stages is itself the result. The argument is that validity, uniqueness, novelty and FCD say almost nothing about medicinal plausibility, so generators that look strong on GuacaMol or MOSES can be producing compounds that no discovery team would carry forward.","tasks":"Generative model evaluation, multi-stage hit-identification filtration","note":"The thresholds at every stage are choices, and different choices reorder the leaderboard — the funnel encodes one company's view of what a plausible hit looks like. Stages five and six inherit every known failure of docking as an affinity proxy, so a generator can still be rewarded for exploiting a scoring function, and PoseCheck-style strain analysis is a useful companion. Preprint from an industry group, not yet peer reviewed or independently reproduced, and no public repository could be confirmed at time of indexing.","added":"2026-08-10"},{"name":"BioXArena","cat":"agent","author":"Loka Li, Duzhen Zhang, Xingbo Du et al. (MBZUAI AI4Bio)","repo":"mbzuai-ai4bio/BioXArena","site":"https://arxiv.org/abs/2605.15766","consortium":"None","desc":"Seventy-six end-to-end BioML coding tasks across nine domains — sequence, single-cell, structure, network biology, chemical biology, perturbation dynamics, phenotype-disease, imaging and text-integrated — where the agent must write runnable code, train a model and submit predictions for a private test split. Each task ships as a public data capsule with hidden labels and a held-out grader, scored with a biology-aware metric rescaled to 0-1. The distinguishing move against BioML-bench is modality realism: most tasks combine several input sources and over half are genuinely multi-modal, mixing tables, images, text, molecular sequences, omics matrices and protein structures rather than handing the agent one clean CSV. Eleven agent configurations were run in a shared two-hour single-GPU sandbox; the best averaged 0.666 and no agent led across all nine domains.","tasks":"Agentic BioML code generation, multi-modal model building, held-out grading","note":"Tasks are curated from public primary sources (BindingDB, Tox21, STRING, KEGG, Reactome, DisGeNET, HAM10000, LIDC-IDRI and similar), so hidden test labels guard against answer leakage but not against a model having seen the dataset — contamination controls are described rather than measured. The fixed two-hour single-GPU budget is a hard constraint that partly measures engineering thrift rather than modelling ability, and rescaling nine domains of very different metrics onto one 0-1 average makes the headline number easier to read than to interpret; read the per-domain table. Coverage explicitly omits spatial omics, whole-slide pathology, raw microscopy time series, flow cytometry, mass spectrometry and longitudinal EHR. The bigger practical caveat is scoring: as of indexing the ground-truth answer files for all 76 tasks are unreleased and the only way to get a score is to email your output directory to the authors, who grade it and decide whether to list you. Contamination-resistant, but it means nobody can independently reproduce or audit the published leaderboard.","added":"2026-08-11"},{"name":"FoldBench","cat":"structure","author":"Sheng Xu, Qiantai Feng, Lifeng Qiao, Shuangjia Zheng & Siqi Sun (BEAM Lab)","repo":"BEAM-Labs/FoldBench","site":"https://www.nature.com/articles/s41467-025-67127-3","consortium":"None","desc":"A unified all-atom benchmark of 1,522 low-homology biological assemblies covering nine tasks: protein-protein (279), antibody-antigen (172), protein-ligand (558), protein-peptide (51), protein-RNA (70) and protein-DNA (330) interfaces plus protein, RNA and DNA monomers. Where PoseBusters and DockGen measure only small-molecule docking, FoldBench scores the whole AlphaFold3-class output space on one homology-controlled target set, so a co-folding model can be compared across modalities rather than cherry-picked on the one it happens to be good at. Success is LRMSD < 2 A with LDDT-PLI > 0.8 for protein-ligand and DockQ >= 0.23 for other interfaces, with a maintained public leaderboard and a submission route for new models.","tasks":"All-atom structure prediction across six interaction types + monomers","note":"The homology cut is defined against PDB data before 2023-01-13, so any model trained past that date — Boltz-2 and RosettaFold3 on the current board are flagged for exactly this — may have seen the targets or close homologs, and those rows are not comparable to the strictly valid ones. Read that caveat before quoting a number. Two findings are worth carrying: protein-ligand accuracy falls off sharply as ligand novelty increases, and antibody-antigen remains unsolved, with over half of all predictions failing even for AlphaFold 3. Reference results are produced by the benchmark authors running the models themselves, which is more trustworthy than self-reported numbers but concentrates every inference choice in one group.","added":"2026-08-12"},{"name":"PXMeter / PXM benchmark datasets","cat":"docking","author":"Wenzhi Ma, Zhenyu Liu, Jincai Yang & Wenzhi Xiao (ByteDance, Protenix team)","repo":"bytedance/PXMeter","site":"https://www.biorxiv.org/content/10.1101/2025.07.17.664878v1","consortium":"None","desc":"An Apache-2.0 evaluation toolkit plus the PXM series of curated datasets, built on the argument that co-folding papers disagree less about models than about scoring conventions. It performs full-atom entity matching, sequence alignment and chain/atom permutation before computing LDDT, DockQ, pocket-aligned ligand RMSD and PoseBusters validity, so two groups scoring the same predictions get the same number. The dataset pipeline for the RecentPDB low-homology set — filtering, homology scans, clustering and subset labelling — is released in full and can be re-run over any custom time window, so the benchmark can be regenerated as the PDB grows rather than freezing in place.","tasks":"Unified co-folding evaluation, dataset curation, protein-ligand and antibody-antigen subsets","note":"Built and maintained by the team behind Protenix, so the evaluation conventions were chosen by a party that also ships a model scored under them — worth knowing even though the code is open and the metrics are standard. The benchmarking workflow ships only in the source repository, not the PyPI package. The authors state development and testing used RCSB PDB CIF files exclusively, so references from other sources are unvalidated. This is infrastructure for fair comparison rather than a leaderboard with a community around it.","added":"2026-08-12"},{"name":"2026ARK-AB","cat":"structure","author":"Aureka Research (released with OpenDDE)","repo":"aurekaresearch/OpenDDE","site":"https://arxiv.org/abs/2607.03787","consortium":"None","desc":"A 164-complex antibody-antigen co-folding test set spanning 159 unique interface clusters, curated from recent low-homology PDB depositions and released Apache-2.0 alongside the OpenDDE model. It exists because the antibody-antigen slices of FoldBench and PXMeter draw on older PDB windows, and antibody-antigen is the one category where every co-folding model still fails more than half the time, so a fresher held-out set has real diagnostic value. Reported top-ranked success rates are OpenDDE 66.4%, OpenFold3 33.1%, Boltz-1 25.0% and Chai-1 16.7%, rising to 80.1% for OpenDDE under oracle selection.","tasks":"Antibody-antigen co-folding, low-homology held-out evaluation","note":"Curated by the authors of the model that tops it and shipped in the same package — the standard vendor-benchmark conflict. The gap between 66.4% top-ranked and 80.1% oracle shows how much of the headline depends on ranking rather than sampling. Baseline numbers for competing models are self-run by Aureka, not submitted by their authors. At 164 complexes the set is small enough that differences of a few percent are noise. Treat it as a supplementary panel alongside FoldBench-AB and PXM-22to25-Antibody rather than a standalone verdict; the useful test is whether an independent group reproduces the ordering.","added":"2026-08-12"},{"name":"HeurekaBench / sc-HeurekaBench","cat":"agent","author":"mlbio-epfl (Maria Brbic lab, EPFL)","repo":"mlbio-epfl/HeurekaBench","site":"https://arxiv.org/abs/2601.01678","consortium":"None","desc":"A framework for manufacturing benchmarks rather than a fixed benchmark: given a published study and its code repository, a semi-automated pipeline uses several LLMs to extract the study's insights and generate candidate analysis workflows, which are then verified against the findings the paper actually reported. The output is open-ended research questions with grounded answers, aimed at the co-scientist setting where a system must analyse experimental data, interpret it and produce an insight rather than hit a prediction metric. Instantiated as sc-HeurekaBench over single-cell biology and used to compare current single-cell agents; adding a critic module lifted open-source agents by up to 22% and closed much of the gap to closed-source ones. ICLR 2026.","tasks":"Open-ended data-driven research, single-cell agent evaluation, benchmark generation","note":"Questions are LLM-generated and validated against what a paper reported, so the ground truth is 'what the original authors concluded' rather than what is true, and it inherits their errors — a real ceiling on open-ended discovery evaluation. Scoring open-ended answers needs a judge, so the numbers carry judge bias. Source studies are published and public, so contamination is plausible and is not measured. Only the single-cell instantiation exists; the claim that the framework generalises to other domains is untested.","added":"2026-08-12"},{"name":"BOOM","cat":"admet","author":"Evan R. Antoniuk, James Diffenderfer, Bhavya Kailkhura et al. (Lawrence Livermore National Laboratory)","repo":"FLASK-LLNL/BOOM","site":"https://arxiv.org/abs/2505.01912","consortium":"None","desc":"Benchmarks for Out-Of-distribution Molecular property predictions: instead of splitting by scaffold, BOOM holds out the extremes of the property distribution itself, so a model is scored on whether it can extrapolate beyond the property range it was trained on. Over 150 model-task combinations were run, including chemical foundation models, GNNs and descriptor baselines.","tasks":"Property-extrapolation OOD prediction","note":"Headline result is that no model generalises: the best average OOD error was still 3x the in-distribution error, and pretrained chemical foundation models were not better than high-inductive-bias models. Caveat for drug discovery use — most of the target properties are computed (QM/simulation-derived) rather than experimental assay endpoints, so it measures chemical extrapolation, not assay transfer.","added":"2026-08-13"},{"name":"PSBench","cat":"structure","author":"Pawan Neupane, Jian Liu & Jianlin Cheng (University of Missouri)","repo":"BioinfoMachineLearning/PSBench","site":"https://arxiv.org/abs/2505.22674","consortium":"None","desc":"A large-scale benchmark for estimation of model accuracy (EMA) on protein complexes: ~1.4M structural models across five labelled datasets, four generated during CASP15 and CASP16 and one curated from PDB entries deposited July 2024 to August 2025. Each model carries global, local and interface-level quality labels, with baseline EMA methods and evaluation metrics supplied.","tasks":"Model quality estimation / ranking for complexes","note":"The model pool is whatever CASP predictors submitted, so the quality distribution reflects CASP-era methods and is skewed toward strong predictors — EMA methods tuned on it may not transfer to models sampled from a single modern co-folding tool. Overlaps with CASP data that some EMA methods already train on; check provenance before believing a leaderboard gain.","added":"2026-08-13"},{"name":"BioMedArena","cat":"agent","author":"Jinge Wu, Hongjian Zhou et al. (UCL / University of Oxford)","repo":"AI-in-Health/BioMedArena","site":"https://arxiv.org/abs/2605.06177","consortium":"None","desc":"An evaluation harness rather than a single benchmark: it decouples six layers of biomedical agent evaluation (benchmark loading, tool exposure, tool selection, harness mode, context management, scoring) and exposes 160+ biomedical benchmarks and 75 tools behind one interface, so agent scaffolds can be compared under a fixed environment. Ships reference harnesses and context-management strategies that can be swapped onto any backbone.","tasks":"Biomedical deep-research agent evaluation","note":"Its value is controlling for scaffold and tool access, which most agent papers conflate with model quality. But it inherits every weakness of the 160+ benchmarks it wraps — a high aggregate score mostly reflects the mix of underlying tasks, so read per-benchmark results, not the average.","added":"2026-08-13"},{"name":"EpiBench","cat":"agent","author":"Zirui Wang, Jiaqi Wang, Tingjun Hou & Odin Zhang (Zhejiang University)","repo":"","site":"https://arxiv.org/abs/2608.06022","consortium":"None","desc":"A closed-book, sequence-only benchmark for epitope reasoning in LLMs: 1,609 samples grounded in structural antibody-antigen contacts, functional B-cell assays and deep mutational scanning escape data, spanning five linked tasks (targetable region discovery, antibody-conditioned epitope identification, epitope binning, functional epitope assessment, escape assessment). Unlike existing epitope resources it evaluates decisions across an antibody development workflow rather than a single isolated prediction task.","tasks":"Epitope reasoning, antibody escape, epitope binning","note":"Closed-book and sequence-only by design, so it measures what a model has internalised rather than what a structural tool can compute — a fair diagnostic for LLMs but not a substitute for structure-based epitope prediction. Nine general-purpose LLMs were evaluated at release and all showed weak long-context residue localisation; no public repository confirmed yet, so treat reproduction as untested.","added":"2026-08-14"},{"name":"JUMP-lite / Nahual","cat":"target","author":"Alán F. Muñoz, Johan Fredin Haslum, Anne E. Carpenter & Shantanu Singh (Broad Institute)","repo":"","site":"https://arxiv.org/abs/2608.07632","consortium":"None","desc":"A 116 GB curated subset of the 115 TB JUMP Cell Painting corpus — roughly 1000x smaller — selected to preserve phenotypic diversity, plus Nahual, a framework that runs each representation model in an isolated environment so methods can be swapped without dependency clashes. Benchmarks five cell-representation methods (CellProfiler, MorphEM, OpenPhenom, SubCell, DINOv2) under standardised phenotypic activity and consistency metrics.","tasks":"Image-based cell representation benchmarking","note":"The subset relies on lossy JPEG XL compression; the authors measure the cost at about 1.2% across eleven retrieval tasks, which is small but not zero and has not been checked for tasks beyond those eleven. Being a subset, absolute numbers are not comparable to full-JUMP results.","added":"2026-08-14"},{"name":"IMPROVE Expanded Drug Response Benchmark","cat":"target","author":"IMPROVE project team (NCI-DOE Collaboration)","repo":"","site":"https://arxiv.org/abs/2608.11444","consortium":"None","desc":"A large expansion of the IMPROVE (Innovative Methodologies and New Data for Predictive Oncology Model Evaluation) benchmark, integrating PharmacoDB and smaller sources into millions of drug response measurements with broader multi-omics coverage and more than 50,000 additional compounds. Keeps IMPROVE's unified data schema and evaluation protocol, including drug-blind and disjoint splits that test generalisation to unseen compounds.","tasks":"Cancer cell line drug response prediction","note":"Cell line response is a weak proxy for clinical efficacy, and aggregating pharmacogenomic sources inherits their inter-lab variability — reported gains are on the expanded data itself, so cross-source consistency has not been independently audited.","added":"2026-08-14"},{"name":"TadA-Bench","cat":"structure","author":"TadA-Bench authors (Shanghai Jiao Tong University)","repo":"","site":"https://arxiv.org/abs/2606.02624","consortium":"None","desc":"A million-variant wet-lab replay benchmark built from 31 rounds of a real TadA directed-evolution campaign. It preserves the campaign chronology and poses a future-round task: given earlier rounds, rank variants that appear only in later rounds, with aligned DNA, RNA and protein views and a label-unification pipeline that reconciles noisy enrichment measurements across rounds.","tasks":"Future-round variant ranking, finite-budget selection","note":"The headline result is the gap between random splits, where protein language models interpolate well, and future-round ranking, where they do much worse — a direct warning about ProteinGym-style interpolation scores. Single protein family, so generality beyond TadA is unproven. Backfilled: published June 2026 but previously missing from this index.","added":"2026-08-14"},{"name":"scBench","cat":"agent","author":"Kenny Workman, Zhen Yang, Harihara Muralidharan, Aidan Abdulali & Hannah Le (LatchBio)","repo":"latchbio/scbench","site":"https://arxiv.org/abs/2602.09063","consortium":"None","desc":"195 verifiable single-cell RNA-seq analysis problems, each pairing an AnnData snapshot with a natural-language task prompt and a deterministic grader that returns pass/fail. Tasks span six sequencing platforms (BD Rhapsody, Chromium, CS Genetics, Illumina, Mission Bio, Parse Bio) and six workflow stages from QC to differential expression, and are constructed so that an agent which answers from prior knowledge without touching the data fails.","tasks":"Agentic scRNA-seq analysis, QC, clustering, cell typing, DE","note":"Only six canonical examples are public; the full 195-eval set is withheld to prevent training contamination, so third parties cannot independently reproduce the leaderboard. Best frontier scores sit around 58%, which leaves plenty of headroom, but the deterministic numeric graders mean a defensible alternative analysis choice can still be marked wrong.","added":"2026-08-16"},{"name":"scBench-Long","cat":"agent","author":"Ian Diks, Zhen Yang, Arjun Banerjee, Tim Proctor & Kenny Workman (LatchBio)","repo":"latchbio/scbench-long","site":"https://arxiv.org/abs/2606.26563","consortium":"None","desc":"The long-horizon counterpart to scBench: 21 evaluations in which an agent starts from raw or near-raw data and must recover the study's actual scientific conclusion with no prescribed method. Study systems include melanoma CD8 T-cell reactivity, CD8 RNA+ATAC regulatory inference, human-monkey chimera development, KRAS-driven lung tumour ageing and lethal COVID-19 lung pathology, spanning paired scRNA/TCR, multiome, cross-species and single-nucleus data.","tasks":"Long-horizon agentic biology, conclusion recovery","note":"Brutally hard — across 1,068 trajectories the best model-harness pair passed 16/63 runs (25.4%). Only four evals are released publicly, so the headline numbers are not independently reproducible, and n=21 tasks over five study systems means per-system variance dominates any ranking.","added":"2026-08-16"},{"name":"MVCBench","cat":"target","author":"Bo Li, Bob Zhang & Qianqian Song (University of Macau / University of Florida)","repo":"QSong-github/MVCBench","site":"https://www.biorxiv.org/content/10.64898/2026.04.22.720110v1","consortium":"None","desc":"A benchmark of representation choice rather than model architecture: 12 drug molecular and 12 gene representation methods evaluated across ~1.1 million drug-induced profiles for predicting both transcriptomic and high-content imaging phenotypes. Splits cover in-distribution and out-of-distribution settings over unseen compounds, cell lines, assay plates and whole datasets.","tasks":"Drug-induced transcriptomic and morphological phenotype prediction","note":"Its most useful finding is an asymmetry: advanced molecular representations help morphology prediction but barely beat classical fingerprints for gene expression. Performance collapses under cross-dataset and cross-platform shift, so in-distribution numbers here should not be read as generalisation. A preprint with no independent reproduction yet.","added":"2026-08-16"},{"name":"DrugDiscoveryBench","cat":"agent","author":"Scale Labs and Phylo","repo":"scaleapi/DrugDiscoveryBench","site":"https://labs.scale.com/leaderboard/drugdiscoverybench","consortium":"None","desc":"Eighty-two expert-written tasks covering the computational work that precedes a molecule: target identification, patent and database research, structural analysis, molecular biology and lead optimisation. Each task is a self-contained Harbor trial run inside a pinned Docker image that ships the full Biomni tool suite plus a baked read-only data lake, so the agent pulls data, runs the analysis and writes an answer without leaving the container. Grading is by LLM judge against expert-authored ground truth plus separate outcome and process rubrics. Its most useful result is a harness ablation: handing agents expert-written analysis plans lifted coverage to 76 of 82 tasks, which locates the failure in deciding what to do rather than in operating the tools.","tasks":"Target ID, patent and database research, structural analysis, lead optimisation, agentic tool use","note":"Rubrics and reference answers are deliberately kept out of the public repo and distributed through a gated Hugging Face dataset, which protects against training contamination but means the grading criteria cannot be inspected before requesting access and the leaderboard is run by the party that holds the answers. LLM-judge scoring against rubrics carries the usual judge bias. Because every task runs inside one Biomni-based image, a score partly measures fluency with that specific toolchain rather than discovery ability in general. Eighty-two tasks with best scores near 50% means a handful of items can move a ranking, and the infrastructure cost is real: roughly 23 GB on disk and up to 45 minutes per task. Brand new and not yet independently reproduced.","added":"2026-08-17"},{"name":"URSA (Utilitarian Retrosynthesis Assessment)","cat":"synthesis","author":"Zagribelnyy, Ilin, Bondarev, Morgunov, Lin, Kuznetsov, Shayakhmetov, Aladinskiy, Aliper & Zhavoronkov (Insilico Medicine)","repo":"","site":"https://arxiv.org/abs/2607.04688","consortium":"Insilico","desc":"Scores whole synthetic routes on two axes that are usually collapsed into one: formal validity, meaning convergence to commercially available starting materials, and chemical plausibility, meaning whether a chemist would actually run the proposed reactions. Conventional end-to-end retrosynthesis engines and LLMs are evaluated on the same set of novel, diverse targets, which is rare — the two families are normally benchmarked apart on incomparable metrics. It is the route-level companion to the same group's single-step ChemCensor work already indexed here.","tasks":"Multi-step route planning, chemical plausibility of routes, LLM vs specialist retrosynthesis comparison","note":"Overlaps the ChemCensor / CREED / URSA-expert-2026 entry in this index — same group, same target set lineage — so treat the two as one programme rather than independent evidence. No public repository could be confirmed at time of indexing, and plausibility scoring ultimately encodes one company's chemists' judgement, so the ranking is only as transferable as that judgement. Preprint, not independently reproduced.","added":"2026-08-17"},{"name":"InteractBind","cat":"affinity","author":"Zhaohan Meng et al.","repo":"","site":"https://arxiv.org/abs/2605.24045","consortium":"None","desc":"About 100k protein-ligand pairs over ~11k proteins and ~9k ligands, each carrying a binary binding label, an affinity value, binding-site annotations and a residue-atom interaction map across six non-covalent interaction types. The evaluation is deliberately two-tier: a model is scored not only on whether it predicts binding but on whether its interaction map lands on the residues that actually do the binding. Affinity and protein-similarity-controlled splits are supplied for generalisation testing. Where PLINDER and PoseBusters test structure-based methods on pose geometry, this tests sequence-based DTI models on whether their attention corresponds to physical reality.","tasks":"Drug-target binding prediction, binding-site localisation, non-covalent interaction typing","note":"The headline finding is the gap it exposes: across eight evaluated sequence-based and interaction-aware models, one reaching 98.3% binary binding accuracy localised the true binding site only 21.6% of the time — strong evidence that these models learn binding likelihood from ligand and protein priors rather than interaction physics. Caveats: labels and interaction maps are derived from structural and assay databases, so annotation quality is inherited rather than measured; the localisation metric rewards agreement with one annotated site and will penalise a model that flags a genuine secondary or allosteric pocket. Distributed as a Hugging Face dataset (Zhaohan-Meng/InteractBind) with no GitHub repository confirmed at time of indexing. NeurIPS 2026 datasets and benchmarks submission, not yet independently reproduced.","added":"2026-08-18"},{"name":"FLAb / FLAb2 (Fitness Landscape for Antibodies)","cat":"structure","author":"Michael Chungyoun & Jeffrey J. Gray (Johns Hopkins)","repo":"Graylab/FLAb","site":"https://r2.graylab.jhu.edu/flab","consortium":"None","desc":"The largest open therapeutic-antibody developability benchmark: 241 curated datasets and over three million antibody assay measurements aggregated from public studies across seven properties – expression, thermostability, immunogenicity, aggregation, polyreactivity, binding affinity and pharmacokinetics. Uniform CSV format (heavy/light sequence plus a fitness column) with zero-shot and few-shot scoring scripts for roughly thirty protein language, inverse-folding, diffusion and biophysical models. Where AbBiBench focuses on affinity maturation against specific antigens, FLAb2 is about the developability properties that actually kill antibody candidates downstream.","tasks":"Antibody developability prediction, zero-shot and few-shot fitness ranking","note":"The benchmark exists mainly to report a negative result, and it is a blunt one: protein AI models produce no statistically significant correlation on roughly 80% of the developability datasets, and no model correlates across all properties or even across multiple datasets of the same property. The ablation is the sharper finding – germline edit distance alone accounts for about 40% of protein language models' apparent predictive power, meaning much of what looks like learned biophysics is evolutionary proximity. Caveats on the data side: assays are aggregated across labs with heterogeneous protocols and units, so cross-dataset comparison is not meaningful, and fifteen datasets with non-standard fitness columns are excluded from automated scoring. bioRxiv preprint (Dec 2025); the repository is a living benchmark and has moved past the numbers quoted in the paper.","added":"2026-08-19"},{"name":"GeneDisco","cat":"target","author":"Arash Mehrjou, Patrick Schwab et al. (GSK.ai, Oxford, MIT)","repo":"genedisco/genedisco","site":"https://arxiv.org/abs/2110.11875","consortium":"None","desc":"A benchmark suite for active learning and experimental design over CRISPR genetic perturbation screens – the question is not how well you predict, but which experiments you should run next given a fixed budget of cycles and batch size. Bundles curated public perturbation datasets, several feature sets (Achilles, STRING, CCLE among them) and reference implementations of standard acquisition functions, with a defined loop protocol so custom acquisition strategies plug in and are scored on the same footing. ICLR 2022, with an accompanying GSK-run challenge.","tasks":"Batch active learning, acquisition function evaluation, target discovery rate","note":"The uncomfortable result that has followed this benchmark since publication is that sophisticated acquisition functions rarely beat random selection by a convincing margin on most of its tasks, and the ranking is unstable across seeds and batch sizes – report multiple seeds and always include the random baseline. The repository has been quiet since 2022 (pinned to Python 3.8, dependencies are ageing), so treat it as a stable historical reference rather than a maintained platform. The underlying screens are cell-line-specific, so 'target discovery rate' here is a proxy several steps removed from a validated drug target.","added":"2026-08-19"},{"name":"CoFold Arena","cat":"structure","author":"Bowen Jing (MIT CSAIL)","repo":"","site":"https://cofoldarena.ai","consortium":"None","desc":"A PDB-synced leaderboard for open co-folding models, refreshed weekly against structures released since September 2024 so the evaluation set grows instead of freezing. Around fifteen model variants – AF-Multimer through OpenDDE – are run by the maintainer on the same antibody-antigen targets, with every prediction viewable and downloadable rather than reported only as a summary score. The distinguishing feature is a dynamic release-window selector: the evaluation set can be rebuilt over any date range, so a model's score can be watched as targets cross its training cutoff, which turns leakage from an assumption into something directly observable. That addresses a problem the static sets here share – FoldBench cuts homology at January 2023 and 2026ARK-AB is a fixed 164 complexes, so both age as models retrain.","tasks":"Antibody-antigen co-folding, rolling PDB-synced evaluation, training-leakage probing","note":"Antibody-antigen only at present; protein-ligand benchmarking is advertised as coming but is not live, so this does not yet cover small-molecule discovery. Maintained by one researcher rather than a consortium, with no paper, no peer review and no public repository confirmed at time of indexing – the evaluation code, metric definitions and inference settings cannot be inspected, and every number comes from one person's runs of models whose authors did not submit them. The rolling window is both the strength and the catch: a score quoted today is not reproducible later unless the date range is pinned alongside it. New and not yet independently reproduced.","added":"2026-08-19"},{"name":"KinDEL","cat":"affinity","author":"Benson Chen, Tomasz Danel et al. (insitro)","repo":"insitro/kindel","site":"https://github.com/insitro/kindel","consortium":"None","desc":"One of the first large public DNA-encoded library datasets: ~81M compounds screened against two kinases (MAPK14 and DDR1), released with on-DNA enrichment counts plus off-DNA biophysical validation (SPR) on a held-out subset. The off-DNA data is what makes it a benchmark rather than a data dump – models are scored on whether enrichment-trained predictions survive a real affinity assay.","tasks":"DEL enrichment modelling, hit identification, off-DNA validation","note":"Two targets from one library and one screening campaign, so chemical diversity is narrow and the validation subset is small; strong correlation on enrichment counts does not transfer to the off-DNA set for most methods. Data lives in an AWS S3 bucket rather than in the repo, which makes fully offline reproduction awkward.","added":"2026-08-20"},{"name":"CA-DEL","cat":"affinity","author":"Mutian He, Hanqun Cao et al.","repo":"","site":"https://arxiv.org/abs/2605.07439","consortium":"None","desc":"A multi-target, multi-modal DEL benchmark built around three homologous carbonic anhydrase isoforms (CAII, CAIX, CAXII), so the task is selectivity between close relatives rather than binder-versus-nonbinder. Pairs noisy sequencing-derived enrichment signal with a validation set of experimentally determined affinities pulled from ChEMBL, and adds structural modality on top of the count data.","tasks":"DEL enrichment modelling, isoform selectivity, affinity validation","note":"The selectivity framing is the contribution; carbonic anhydrase is also one of the most over-represented targets in public affinity data, so ChEMBL-derived validation risks overlapping whatever a pretrained model already saw. Repository and Zenodo deposit were announced as forthcoming in the paper – link is to the arXiv record until the code is confirmed live.","added":"2026-08-20"},{"name":"Ab-VS-Bench","cat":"agent","author":"Ab-VS authors (bioRxiv 2025.07.26.666985)","repo":"","site":"https://www.biorxiv.org/content/10.1101/2025.07.26.666985v1.full","consortium":"None","desc":"Ports the small-molecule virtual screening workflow to antibodies and to language models: scoring (predict antibody-antigen binding affinity), ranking (order antibodies by affinity or thermostability) and screening (recover high-affinity binders from a library), all posed as natural-language instructions. Compares zero-shot LLMs, instruction-tuned LLMs and multimodal models given antibody-specific embeddings.","tasks":"Antibody affinity scoring, ranking, library screening","note":"Reformatting existing experimental antibody datasets into instruction format means public sequences and measurements may already sit in the LLM pretraining corpus, and the paper does not fully resolve that contamination question. Useful mainly as a demonstration that sequence-only LLMs remain far behind dedicated antibody models – treat it as a floor test, not a ranking of therapeutic design methods.","added":"2026-08-20"},{"name":"RetroCast / Procrustean Bed","cat":"synthesis","author":"Anton Morgunov & Victor S. Batista (Yale)","repo":"ischemist/project-procrustes","site":"https://arxiv.org/abs/2512.07079","consortium":"None","desc":"An evaluation harness that decouples retrosynthesis prediction from scoring. Air-gapped adapters cast the mutually incompatible outputs of AiZynthFinder, Retro*, DirectMultiStep, SynPlanner, Syntheseus, ASKCOS, RetroChimera, DreamRetro, MultiStepTTL, SynLlama and PaRoutes into one Pydantic route schema, so models can be compared without writing a bespoke parser per paper. It ships two new benchmark series that fix documented defects in PaRoutes n5: a Reference series (ref-lin-600, ref-cnv-400, ref-lng-84) stratified by route length and topology because 74% of PaRoutes routes are 3-4 steps and mask failures on hard targets, and a Market series (mkt-lin-500, mkt-cnv-160) scored against a 300k-compound catalogue of buyables under $100/g rather than made-to-order stock. Scoring is bootstrapped with 95% CIs, pairwise tournaments and probabilistic ranking; SynthArena provides route-level visual diffing.","tasks":"Multi-step retrosynthesis, route solvability, cross-model standardisation, statistical ranking","note":"The benchmark sets are resampled from PaRoutes, so they inherit the USPTO patent-reaction distribution and its bias toward reactions companies chose to patent – the stratification fixes the skew, not the source. 'Solvability against a buyables catalogue' is still a paper metric: no route here has been run in a lab, and catalogue availability drifts, so scores are only reproducible against the frozen stock files. The scored-prediction database standardises other groups' published outputs, which means comparisons depend on how each model was configured by these authors rather than by its own developers. The canonical repo is ischemist/project-procrustes; batistagroup/retrocast is a mirror and the pip package is 'retrocast'. Preprint, not yet independently reproduced.","added":"2026-08-21"},{"name":"MolGenBench","cat":"generative","author":"Duanhua Cao, Zhehuan Fan, Jie Yu, Mingan Chen et al.; Mingyue Zheng (Shanghai Institute of Materia Medica)","repo":"","site":"https://www.biorxiv.org/content/10.1101/2025.11.03.686215v1","consortium":"None","desc":"Evaluates 17 molecular generative models against 120 protein targets using 5,433 chemical series and 220,005 experimentally confirmed actives, and adds a hit-to-lead optimisation task alongside conventional de novo generation – the stage where most generative benchmarks simply stop. The metric worth borrowing is TAScore, which compares a model's success rate conditioned on a specific target against its background success rate across all targets. That isolates whether output is actually driven by target information or is just a general bias toward active-looking chemistry, a confound that validity/uniqueness/novelty triples cannot detect.","tasks":"De novo generation, hit-to-lead optimisation, target awareness, recovery of known actives","note":"Success is defined as recovering known actives, which rewards rediscovery and structurally penalises a model that proposes a genuinely novel but valid chemotype – the opposite of what generative design is for. The 120 targets are those with enough public activity data to build series from, so coverage skews to kinases and other well-mined families and says little about novel or hard targets. Most of the 17 models were trained on overlapping public actives, so contamination against the evaluation set is plausible and is not quantified in the preprint. Code was announced under a GitHub path that is truncated in the available sources, so no repository is listed here rather than risk a wrong one – check the paper's data availability statement. bioRxiv preprint, not independently reproduced.","added":"2026-08-21"},{"name":"DisProtBench","cat":"structure","author":"Xinyue Zeng, Tuo Wang, Adithya Kulkarni, Alexander Lu, Alexandra Ni, Phoebe Xing, Junhan Zhao, Siwei Chen, Dawei Zhou","repo":"Susan571/DisProtBench","site":"https://arxiv.org/abs/2507.02883","consortium":"None","desc":"An IDR-centric benchmark for protein structure prediction, covering disease-relevant intrinsically disordered regions, GPCR–ligand pairs and multimeric complexes – the cases that CASP-style and FoldBench-style evaluations largely skip in favour of well-folded domains. Its contribution is Functional Uncertainty Sensitivity (FUS), a metric that stratifies downstream task performance by the model's own confidence rather than reporting a single global accuracy number. The headline finding is that protein–protein interaction prediction collapses in disordered regions while structure-based drug discovery holds up comparatively well, a task-dependent split that global metrics such as TM-score and lDDT hide entirely.","tasks":"Disordered-region structure prediction, protein-protein interaction, GPCR-ligand modelling, multimeric complex prediction, uncertainty-stratified functional evaluation","note":"Ground truth for disordered regions is the core problem the benchmark is trying to measure and it cannot fully escape it – IDRs are conformational ensembles, so any single reference structure is a convenience rather than a truth, and the curated labels inherit whatever bias the source databases carry. The GPCR-ligand and multimer subsets are drawn from public structures, so contamination against models trained on the PDB is plausible and is not quantified. FUS depends on models reporting calibrated confidence, which they do to very different standards, making cross-model FUS comparisons less clean than the paper implies. Backfill – v1 dates to June 2025 with a substantially revised v2 in February 2026; not independently reproduced.","added":"2026-08-23"},{"name":"PyTrial","cat":"clinical","author":"Zifeng Wang et al.; Jimeng Sun (University of Illinois Urbana-Champaign)","repo":"RyanWangZf/PyTrial","site":"https://arxiv.org/abs/2306.04018","consortium":"None","desc":"A benchmark and reference-implementation suite covering 34 ML algorithms across six clinical trial tasks: patient outcome prediction, trial site selection, trial outcome prediction, patient-trial matching, trial similarity search and synthetic data generation. It ships 23 ML-ready datasets with a uniform four-step load/specify/train/evaluate API, which is what separates it from the single-task clinical benchmarks already in this index. Trial site selection and patient-trial matching in particular have almost no other standardised evaluation anywhere.","tasks":"Patient outcome prediction, trial site selection, trial outcome prediction, patient-trial matching, trial similarity search, synthetic patient data generation","note":"The trial-outcome task overlaps heavily with HINT/TOP, which is the source of that split, so treat the two as one result rather than two independent confirmations. Several patient-level datasets are not actually redistributable and require credentialed access (MIMIC, project-specific trial data), so the '23 ML-ready datasets' figure overstates what a new user can run out of the box. Baselines were frozen at 2023 and the repository has seen limited maintenance since, so reported numbers predate current foundation-model approaches. Backfill of an existing gap, not a new release.","added":"2026-08-23"},{"name":"Bento","cat":"docking","author":"Marina A. Pak, Daria Frolova, Sergei A. Nikolenko, Dmitry N. Ivankov et al. (Ligand Pro / Skoltech)","repo":"LigandPro/Bento","site":"https://www.biorxiv.org/content/10.64898/2025.12.30.696741v1","consortium":"None","desc":"A pocket-aware benchmark that runs eleven protein-ligand structure prediction tools – classical docking, deep-learning docking and co-folding – across four test sets and many derived subsets, stratified by pocket structural similarity and by ligand complexity. Unlike PoseBusters and DockGen, which mainly probe generalisation, Bento asks the separate question of how these tools behave specifically on drug-design-relevant chemistry, and reports each class of complex separately rather than collapsing to one success rate. Its headline results: classical and DL docking are roughly equivalent on drug-like ligands (with physics far faster), co-folding wins on structurally complex ligands, and every method degrades on unseen pockets, DL models worst.","tasks":"Pose prediction, cross-method docking comparison, pocket-generalisation stratification","note":"Run by a single commercial group (Ligand Pro) rather than a neutral consortium, and it benchmarks the authors' competitors – read the tool configurations before quoting the numbers, since docking results are notoriously sensitive to preparation and search-exhaustiveness settings. The pocket-aware setup means results are not comparable to blind-docking numbers from PoseBusters or DockGen. Overlaps substantially with PoseBench and PoseX in the tools covered; treat the three as correlated rather than independent confirmations. Posted December 2025 and not yet peer-reviewed or independently reproduced.","added":"2026-08-24"},{"name":"LP-PDBBind (Leak Proof PDBBind)","cat":"affinity","author":"Jie Li, Xingyi Guan, Oufan Zhang, Kunyang Sun, Yingze Wang, Dorian Bagni & Teresa Head-Gordon (UC Berkeley)","repo":"THGLab/LP-PDBBind","site":"https://arxiv.org/abs/2308.09639","consortium":"None","desc":"A reorganised split of PDBBind that clusters by protein sequence, ligand and structural similarity to remove train/test leakage, exposing three progressively stricter clean levels (CL1–CL3) plus covalent flags. Ships with BDB2020+, an independent held-out set built by matching high-quality BindingDB free energies to PDB complexes deposited since 2020 – a genuine time-split rather than a re-shuffle. Retrained AutoDock Vina, IGN, RFScore and DeepDTA baselines are provided so the size of the leakage effect can be read off directly.","tasks":"Binding affinity regression under leakage-controlled and time-split evaluation","note":"This is the reference point for the criticism recorded against PDBbind and CASF-2016 elsewhere in this index: several published scorers lose roughly 0.2 Pearson R when moved onto the no-leak tiers. Caveats: it inherits every curation error and affinity-type inconsistency (Kd/Ki/IC50 mixed) already in PDBBind 2020, and it is frozen at that version, so it does not track the newer commercially licensed releases. Similarity clustering is threshold-dependent, and the CL1–CL3 tiers are a judgement call rather than a physical boundary. Backfill of a long-standing gap – originally 2023, but still the standard leakage-controlled affinity split and not previously indexed here.","added":"2026-08-24"},{"name":"AbRank","cat":"affinity","author":"Chunan Liu, Aurélien Pelissier et al.","repo":"biochunan/AbRank-WALLE-Affinity","site":"https://arxiv.org/abs/2506.17857","consortium":"None","desc":"A large antibody-antigen affinity benchmark that reframes the task as pairwise ranking rather than regression, aggregating over 380,000 binding assays from nine heterogeneous sources. Splits are constructed to increase distribution shift in stages, from point mutations through to entirely unseen antigens and antibodies, and an m-confident filter discards comparisons where the measured affinity gap is smaller than the assay noise. The WALLE-Affinity graph plus protein-language-model baseline is released alongside, and the dataset is mirrored on Kaggle.","tasks":"Antibody-antigen affinity ranking, generalisation to unseen antigens and antibodies","note":"Ranking sidesteps the cross-source calibration problem rather than solving it – assays pooled from nine sources under different conditions are not on a common scale, which is precisely why regression was abandoned, so absolute affinity claims cannot be made from this benchmark. The linked repository is the baseline model implementation rather than a standalone benchmark package, so expect to assemble the evaluation yourself. Overlaps with AbBiBench in underlying sources (SAbDab and mutation scans), so the two should not be treated as independent. Not yet independently reproduced.","added":"2026-08-24"},{"name":"AMPBench-MT","cat":"admet","author":"Ziheng Zhou et al.","repo":"","site":"https://arxiv.org/abs/2607.25518","consortium":"None","desc":"A homology-controlled benchmark for antimicrobial peptides that places binary AMP recognition, species-conditioned pMIC regression, activity spectrum and safety-facing endpoints (haemolysis, cytotoxicity) inside a single protocol with provenance preserved per record. The controlled variable is sequence homology: near-duplicate peptides are prevented from straddling the split, which is the failure mode that inflates almost all published AMP classifier accuracy. Distributed as a Hugging Face dataset; complements PepBenchmark, which is broader across peptide task types but does not condition potency on target species.","tasks":"AMP recognition, species-conditioned potency (pMIC) regression, activity spectrum, haemolysis and cytotoxicity proxies","note":"The authors are explicit that residual source, species and publication overlap survives homology control and report it as a benchmark boundary rather than claiming clean external validation – so this is less leak-proof than the framing suggests. Its own central finding is a warning: across 161 endpoint evaluations, strong binary AMP classification does not predict assay-endpoint behaviour, so do not read recognition AUC as evidence of potency modelling. MIC values pooled across labs carry well-known inter-laboratory variation of a two-fold dilution or more, which sets a hard floor on achievable regression error. Preprint from July 2026, no independent reproduction.","added":"2026-08-24"},{"name":"OpenBind-0","cat":"structure","author":"OpenBind Consortium / AQ Laboratory","repo":"OpenBind-Consortium/OpenBind-0-model-release-info","site":"https://openbind.uk/news/blog-openbind-0-advancing-open-molecular-structure-prediction","consortium":"ASAP","desc":"OpenBind-0 (OB0) is the first fully open-source co-folding model built on OpenFold3 and specialised for predicting protein–small-molecule complex structures, released alongside 717 newly determined ligand-bound structures (547 fragment and 170 hit binding events; 298 unique fragments, 101 unique hits) across FatA, Dengue/Zika RdRp, and EV-A71 2A. The dataset defines discovery-relevant docking and co-folding benchmarks over fragment-to-hit progression. Structures were contributed via the ASAP Discovery Consortium (RdRp) and a Syngenta-funded project (FatA).","tasks":"Protein–ligand co-folding / complex structure prediction; ligand pose accuracy on fragment-to-hit campaign data; evaluation of target-specific fine-tuning gains.","note":"Performance varies dramatically by target — from high accuracy on EV-A71 2A protease down to <10% success on both RdRp systems — so aggregate scores are misleading; the RdRp and FatA sets are hard, small, and target-specific rather than a broad generalization test. Fine-tuning gains are inconsistent across systems, and the model checkpoint uses a near-final (not final) OpenFold3 architecture trained on PDB through June 2025, so benchmarks risk overlap with contemporaneous training data.","added":"2026-08-27"},{"name":"VIALS (Visual Interpretation of Artifacts in the Life Sciences)","cat":"agent","author":"Elaine Lau, Thanuka Udumulla, Lee Izhaki-Tavor, Francisco Guzmán, Nicholas Magazine, Jonas Mueller","repo":"","site":"https://arxiv.org/abs/2608.21357","consortium":"None","desc":"VIALS is a visual question-answering benchmark of 161 interpretation tasks over real-world biotech artifacts — gel blots, microscopy images, plasmid maps, flow cytometry plots, and molecular structures — as encountered in experimental workflows rather than polished publication figures. It probes whether vision-language models can perform domain-specific scientific visual reasoning, finding frontier VLMs perform poorly while domain-expert scientists find the tasks straightforward.","tasks":"Visual QA / image interpretation of life-science experimental artifacts (gels, microscopy, plasmid maps, flow cytometry, molecular structures).","note":"At only 161 tasks the benchmark is small, so results are sensitive to per-item noise and offer limited statistical power for fine-grained model ranking. The human 'straightforward' baseline comes from domain experts and may overstate the gap versus non-expert humans; task-authoring by a small team also risks stylistic or answer-format idiosyncrasies. No public code/data repository is linked on the arXiv listing.","added":"2026-08-27"},{"name":"Chem World","cat":"admet","author":"Tianyou Bai, Huan Wang, Mingchen Gao, Fangyue Lin, Pinze Ren, Zhenlin Zhao, Siming Dong","repo":"","site":"https://arxiv.org/abs/2607.28079","consortium":"None","desc":"Chem World is a large-scale chemical property prediction benchmark integrating 17 diverse datasets with over 800,000 molecular samples, covering properties such as density, electrical conductivity, solubility, and other molecular characteristics under unified evaluation protocols. It is released together with Mixture-PINN, a physics-informed neural network framework the authors propose as a baseline. The benchmark aims to enable fair, standardized comparison of ML methods for molecular and drug-relevant property prediction.","tasks":"Regression/classification for molecular property prediction across 17 datasets (density, electrical conductivity, solubility, and related endpoints).","note":"Coverage is heavily weighted toward bulk physicochemical/materials properties (density, conductivity) rather than ADME/tox endpoints, so its relevance to drug discovery is partial despite the admet tag. The paper co-introduces the Mixture-PINN method, so reported baselines may favour the authors' framework; no public code/data repository is linked on the arXiv listing, and the 'unified protocol' claim cannot be independently verified without a release. Aggregating 17 heterogeneous datasets risks inconsistent splits and units across tasks.","added":"2026-08-27"}]}