AdaptiveFlow bins 69B REAL Space molecules on an 18-D grid, cutting exhaustive docking 5,000-fold

2026-09-05

AdaptiveFlow grids 68.7B Enamine molecules in 18 property bins. A 12M prescreen localizes tranches; FSP1 gave a 283 nM hit from 22M dockings, PARP1 an 8.8 nM inhibitor.

What problem this solves

A typical experimental HTS campaign tests a few hundred thousand compounds. Estimates of synthesizable drug-like space sit near 10^60. Virtual screening already reaches hundreds of millions to billions of molecules, and the bill for exhaustive docking grows with the library. Ready-to-dock public sets such as the 2018 REAL Database and ZINC20 held about 1.5 billion molecules. Enamine's on-demand catalogs are now quoted in the trillions. Docking every compound is no longer a plan.

Docking codes also disagree on file formats, samplers, and scoring functions, and many stacks still ignore GPUs and ARM. The same group's 2020 VirtualFlow platform topped out at a demonstrated 160,000 vCPUs. What was missing is an open stack that can index an ultralarge library, plug in more than a thousand docking protocols, run on cloud hardware, and feed machine-learning loops.

Method

AdaptiveFlow has three modules: AFLP for ligand preparation, AFVS for screening, and AFU as the combined interface. The actual invention is a chemical-space index, not a new scoring function.

The 2022q1-2 Enamine REAL Space starts as 31.5 billion enumerated SMILES. After stereoisomer and tautomer expansion, protonation, and 3D conformer generation it becomes 68.7 billion ready-to-dock molecules. 99% pass Lipinski's Rule of Five; unique Murcko scaffolds number about 1.1 billion. Each molecule gets 28 computed properties. Eighteen of them (molecular weight, logP, H-bond donors and acceptors, rotatable bonds, TPSA, logS, aromatic fraction, formal charge, sp3 carbon fraction, chiral centers, and similar) define a grid. The theoretical cell count is 24.4 billion; about 12 million cells, called tranches, are occupied, averaging about 5,600 molecules each.

Adaptive Target-Guided Virtual Screening (ATG-VS) then works as follows:

A docking protocol is a sampler paired with a scoring function. About 40 programs combine into roughly 1,500 protocols, including DiffDock, TANKBind, EquiBind, AutoDock-GPU, and Vina-GPU. The library is also shipped as SELFIES for generative models. On AWS Batch with spot instances the pipeline scales near-linearly to 5.6 million vCPUs; ligand preparation lost under 0.1% of CPU hours to preemption. Rebuilding QuickVina 2 with the Intel compiler added an 18% speedup.

Results

On a 5-million-compound subset and ten targets (kinases, phosphatases, GPCRs, protein-protein interfaces), ATG with 100,000 total dockings, with or without the MLP, matched or beat a 1-million random screen on mean top-50 docking scores. For KRAS-G12V the means were −11.83 (1M random), −11.95 (100K ATG), and −12.11 kcal/mol (100K ATG plus active learning). For CB1 they were −13.93, −14.14, and −14.25. Pockets that mix large hydrophobic surfaces with polar subpockets gained more from the MLP. Production runs on the full 68.7 billion library, without active learning, still matched or beat a 100-million random ULVS on top scores at far fewer evaluations.

Two wet targets:

TargetDocking budgetMade and testedKey number
FSP112M prescreen + 10M primary42 ordered, 33 madeafi-FSP1-1 Ki 0.283 μM, Kd 98 nM
PARP1prescreen then 100M primary160 made7 inhibitors; iParp1 IC50 8.8 nM

Of the 33 FSP1 compounds, two inhibited more than 50% at 10 μM and shifted melting temperature by at least 1.5 °C. Co-crystal structures put the inhibitors in the coenzyme Q site, matching the docking pose. Cellular thermal shifts of 2.8 °C and 3.1 °C confirmed engagement of human FSP1 for afi-FSP1-1 and afi-FSP1-2. A follow-up screen of 20 analogs of afi-FSP1-1 added four more inhibitors with Ki from the mid-nanomolar to low-micromolar range. The group's 2020 KEAP1-NRF2 campaign needed an unbiased 1.3 billion dockings for nanomolar hits; FSP1 needed 22 million.

iParp1's enzymatic IC50 of 8.8 nM sits near the FDA drug olaparib. The crystal contains a hydrolyzed form, and cell activity suffered from poor membrane permeability.

Why it matters

For people who run virtual screens this is an engineering upgrade of VirtualFlow plus a chemical-space index. A 69-billion ready-to-dock library, SELFIES, 1,500 protocols, and a GPL-2.0 codebase are infrastructure. The ATG claim is concrete: on FSP1, 22 million dockings produced crystallizable nanomolar inhibitors. Each 3D format compresses to about 50 TB and lives on AWS Open Data, so groups do not have to enumerate the space themselves.

Better docking scores are not the same as a higher wet hit rate. The production comparison reports docking scores only. There is no head-to-head wet test of ATG versus random ULVS on the same target. Active learning never ran on the full 69-billion set.

Limitations

The authors say chemical-space focusing is target dependent; feature-poor pockets gain less. Production benchmarks turned active learning off. The MLP is supervised by docking scores, and those scores already correlate weakly with true affinity, so the second filter trains on a noisy label.

FSP1 crystals used an ancestrally reconstructed tetrapod protein with 72.5% identity to human FSP1. CETSA was done in HEK293 cells expressing the human enzyme, so the structure-to-human-activity jump still has a gap. The strongest PARP1 hit hydrolyzes and permeates poorly. The library is the 2022 REAL Space; the paper already notes purchasable catalogs in the trillions, so the grid will age. Enamine employees are co-authors. ChemAxon JChem remains the default ligand-prep path; open-source fallbacks were added later.

Terms

Source

What people are saying

All paper explainers