Energy-Based Transformers learn to think via unsupervised energy minimization, scaling 35% faster than Transformer++
rbhar90 · x · 2026-09-06
An arXiv paper proposes Energy-Based Transformers (EBTs), a new class of Energy-Based Models that learn System 2-style thinking purely from unsupervised pretraining. EBTs assign an energy score to each input-candidate prediction pair, then reframe prediction as gradient descent-based energy minimization until convergence. Unlike existing inference-time compute methods, they need no verifiers or verifiable rewards and generalize across discrete (text) and continuous (visual) modalities. Across training, EBTs scale up to 35% faster than Transformer++ in terms of data, batch size, parameters, FLOPs, and depth.
More from Research
- AI decodes whale 'talk' by probing the latent space of their calls — maier_ak · 2026-09-07
- New paper explores Conformal Prediction for offensive security attacks — chaumian · 2026-09-07
- HiSfM: scaffold-anchored hierarchical SfM tames repeated-structure ambiguity and cuts runtime — zhenjun_zhao · 2026-09-07
- BLASt3R (ECCV'26): uncalibrated bundle adjustment beats all prior calibrated VSLAM methods — zhenjun_zhao · 2026-09-07
- 4-month-old Chinese startup Atomelody debuts Melo-1, beating AlphaFold3 on most benchmarks — 新智元 · 2026-09-07
- Economists keep getting overparameterization wrong: SGD's implicit regularization is the point — Afinetheorem · 2026-09-07