Energy-Based Transformers learn to think via unsupervised energy minimization, scaling 35% faster than Transformer++
rbhar90 · x · 2026-09-06
An arXiv paper proposes Energy-Based Transformers (EBTs), a new class of Energy-Based Models that learn System 2-style thinking purely from unsupervised pretraining. EBTs assign an energy score to each input-candidate prediction pair, then reframe prediction as gradient descent-based energy minimization until convergence. Unlike existing inference-time compute methods, they need no verifiers or verifiable rewards and generalize across discrete (text) and continuous (visual) modalities. Across training, EBTs scale up to 35% faster than Transformer++ in terms of data, batch size, parameters, FLOPs, and depth.
More from Research
- AdaptVPR: route-aware hard positive generation boosts robust visual place recognition — Shunpeng Chen · 2026-09-07
- SC Asia 2027 opens call for papers on supercomputing and AI infra, due Oct 7, 2026 — thoefler · 2026-09-07
- Frontier models double as RL teachers for smaller siblings, argues poster — haider1 · 2026-09-07
- Ex-Google Brain researcher: key algorithmic wins were found under 1e20 FLOPs, then scaled to 1e25 — _arohan_ · 2026-09-07
- GlossoGen paper: LLM agents evolve emergent languages humans can't understand — abenitezburraco · 2026-09-07
- New paper resolves three open problems in online fair division with impossibility results — chaumian · 2026-09-07