Learned latent protein languages cut structure-prediction perplexity 34% and run ~1000x faster than AlphaFold2
Mahdi Pourmirzaei · hf · 2026-10-07
New research introduces two learned latent protein languages that make autoregressive transformers far stronger at protein sequence and structure generation.
Methods:
- PLL: maps sequences to a 4,096-state contextual alphabet (one token per residue) via a frozen ESM-2 encoder
- SLL: adapts GCP-VQVAE Lite with auxiliary sequence/confidence supervision while decoding back to backbone coordinates
- Separate autoregressive pretraining yields PLLM and SLLM
Results:
- PLLM fits a compute-scaling exponent of 0.038 vs 0.020 for amino-acid tokens
- 54% relative reduction in low-entropy samples during unconditional generation
- Swapping in SLL cuts best validation perplexity by 34% for sequence-to-structure prediction
- Latent-token sampling is 1000x faster than MSA-based AlphaFold2 on long proteins
- SLLM's internal token confidence shows early promise for improving inference-time sampling
The authors position learned latent languages as a scalable substrate for autoregressive protein generation.
More from Research
- Mathematician: norms for LLM use in serious math are still unsettled — littmath · 2026-10-07
- IdeaLens debuts: a detector that separates human ideas from AI-written prose — MohitIyyer · 2026-10-07
- SC4AI'27 workshop on social choice for AI alignment at AAAI, deadline Nov 20 — conitzer · 2026-10-07
- Grafting mid-training weights onto RLed models works, challenging the PSM pipeline — repligate · 2026-10-07
- WonderSearch 1.1 claims #1 on all 7 sparse retrieval benchmarks, no vector DB needed — Scobleizer · 2026-10-07
- NIH Pioneer Award grants ~$6M for LLM agents generating cancer hypotheses tested by robots — anshulkundaje · 2026-10-07