2026-08-11
Missing NMR residue assignments signal millisecond motion; Dyna-1, built on an ESM-3 layer, predicts it and captures catalytic and ligand-binding dynamics.
Proteins are not rigid sculptures; they work by interconverting between conformations on the microsecond-to-millisecond (us-ms) timescale. Enzyme catalysis, ligand binding, and signaling all depend on this motion, and it is hard to measure. Nuclear magnetic resonance (NMR) is one of the few techniques that can quantify us-ms dynamics, but it is slow, scattered, and lacks a shared benchmark; the entire field has only on the order of a hundred usable relaxation datasets. AlphaFold made static structure nearly free, while how proteins move still mostly means laborious manual experiments.
The clue is in what NMR data does not contain. In the BMRB database of roughly ten thousand proteins' chemical shifts, many residues simply have no assignment; they produce no detectable NMR signal. This is usually dismissed as missing data. The authors make a bold call: those residues are exchange-broadened by us-ms motion, which smears their signal out, so the gap in the record is the fingerprint of the motion.
Two steps:
The best model, Dyna-1, takes its features from an intermediate layer of ESM-3, EvolutionaryScale's multimodal protein model that jointly encodes sequence, structure, and function. A pretrained language-model layer carries evolutionary and structural priors for free. The authors unify more than a hundred NMR relaxation datasets with the BMRB chemical-shift archive into one resource, the substrate that makes the model possible.
The abstract states qualitative conclusions; the quantitative figures (correlation, AUC, baselines) sit behind the paywall and were not retrieved.
| Item | Finding |
| Training target | Predict whether a residue is missing a chemical-shift assignment |
| Indirect validation | Predictions match us-ms exchange measured by NMR relaxation |
| Best cases | Motions tied to function: catalytic sites, ligand-binding sites |
| Corroboration | Residues in us-ms exchange are more evolutionarily conserved |
The tell is the logic chain, not a single accuracy number. A model trained only to guess which residue is blank in the spectrum lands on predictions that match real measured motion, and is sharpest exactly at catalytic and binding sites where the protein does its work. That transfer from a proxy label to an experimental observable suggests the missing-signal signal captures dynamics, not a statistical artifact.
For structure and drug-design groups, Dyna-1 moves a task that used to demand dedicated spectrometer time down to the cost of a sequence. Catalytic residues and binding pockets are precisely where drugs act and where motion matters most, and where prediction has been hardest. It also templates a reusable trick: missingness as a label, and language-model representations as a bridge to scarce experimental quantities, likely transferable to other measurements that are expensive and full of gaps.
AlphaFold settled static structure; how a structure moves is now the largest open question in structural biology. Dyna-1 does not close it, but it offers a scalable path that plugs into existing protein models.