Zatom-2 unifies molecule, material, and protein generation; low-data designability 67.8%→74.8%

Zatom-2: Multitask Pretraining on Atomistic Data for Generative Modeling across Domains

Miruna Cretu, Alex Abrudan, Antonia Panescu, Tynan Perez, Rishabh Anand, N. Benjamin Erichson, Michael W. Mahoney, Samuel Blau, Joseph Jacobson, Rafael Gómez-Bombarelli, Rex Ying, Tuomas Knowles, Pietro Liò, Alex Morehead

cs.LG, cs.AI

2026-10-08

One atomistic model, Zatom-2, multitask-pretrains on ~5M DFT structures to generate molecules, materials and proteins; low-data protein designability rises 67.8%→74.8%.

What problem this solves

Generative models over 3D atoms are siloed by discipline. Molecule generators (EDM, TABASCO), crystal generators (CDVAE, DiffCSP), and protein generators (RFdiffusion3, BoltzGen) each ship their own architecture and input representation, and none trains one model across molecules, periodic materials, and biomolecules. The split has no physical basis: molecules, materials, and proteins are all collections of interacting atoms, and the local environments that determine structure recur across system types. The data excuse is gone too. OMol25 holds over 100 million high-fidelity DFT calculations spanning organic molecules, metal complexes, electrolytes, and transition states; OMat24 holds roughly 118 million over inorganic compositions; both carry energy and force labels. Zatom-2 asks how such corpora should be used to learn representations that transfer across tasks and domains.

Method

The organizing idea comes from the protein side: generate only the coordinates and read atom identity out of the learned geometry, with no separate generative process over discrete types. Two tokenization schemes share one backbone: atom1 for molecules and materials (one atom per token) and atom14 for proteins (each residue occupies 14 fixed coordinate slots, residue identity recovered by a sequence head).

Results

BenchmarkMetricZatom-2Baselines
MP20 (unrelaxed)MetaSUN yield4.92%Crystalformer 3.1%, Zatom-1-L 0.64%
GEOM-Drugsvalidity96.90%TABASCO 97.6%, Zatom-1 93.6%
GEOM-DrugsPB-valid91.47%TABASCO 91.6%, Zatom-1 94.1%

QM9 validity lands at 95.72%, described as at or near state of the art, and distribution fidelity on OMol25 clearly beats Zatom-1 on AFD and mean force norms. Force conditioning separates low- from high-force regimes with AUROC between 0.76 and 1.00 across three settings and force MAE of 0.69 to 0.93 eV/Å: reliable regime separation, weaker in-regime precision.

The protein transfer experiment carries the most information. Every model is trained or finetuned on SCOPe-2k, just 2,000 protein domains, and evaluated on length extrapolation (129 to 256 residues):

InitializationDesignabilityPer-seq successNovelty
From scratch67.81%34.68%88.81%
Gen + structure pretraining (no force cond.)62.96%33.91%94.63%
Gen + structure + force/energy pretraining74.80%45.79%95.17%
RFdiffusion3 (same 2k data)54.04%28.37%90.44%

Full generative-predictive pretraining buys 7 points of designability and 11 points of per-sequence success over scratch. Pretraining without force conditioning lands below scratch (62.96% vs 67.81%), so the benefit comes from force-aware pretraining, not pretraining per se. In-distribution (50 to 128 residues) RFdiffusion3 still leads 91.07% to 89.68%; Zatom-2 wins specifically in low-data and extrapolation settings. Scaling behaves as expected: full datasets beat 500k-per-domain subsets at equal epochs, parameter scaling adds more, most visibly on OMat24 structure prediction.

Why it matters

The recipe matters more than the leaderboard. This is one atomistic model covering molecule generation, material generation, structure prediction, and energy/force prediction in a single pretraining run, and it shows quantum-chemistry pretraining transferring to protein generation where structural data is scarce. For molecule and material work, a plain Transformer without enforced equivariance matches or beats specialized architectures, and unrelaxed MP20 metastable yield jumps from 3.1% to 4.92%. For protein design, learning local geometry on small molecules first and then finetuning on a few thousand domains is a usable recipe for data-poor settings. Code, data loading, and training configs are released.

Limitations

Author-stated: no exact rotational equivariance (data augmentation instead), no energy-gradient consistency between the energy and force heads, MLIP accuracy below specialized potentials, and in-regime force control remains hard. Pretraining used the 4M OMol25 slice out of 100M+ available, and the scaling results are themselves an argument for larger runs.

Reasons for caution: the protein experiments cover a narrow regime, 2,000 SCOPe domains of at most 128 residues evaluated up to 256, far from antibody or enzyme design. The RFdiffusion3 baseline was trained on the same 2k data, so this is not a comparison against its fully pretrained release, and in-distribution Zatom-2 loses. Force-conditioning AUROC drops to 0.757 on ANI2x, so controllability is uneven across domains. The paper's AI use statement discloses extensive generative-AI assistance in drafting text, code, and experiment parameters; treat the released code as the source of truth for reproduction.

Terms

Source

What people are saying

Related papers

All paper explainers