Energy-Based Transformers learn to think via unsupervised energy minimization, scaling 35% faster than Transformer++

rbhar90 · x · 2026-09-06

An arXiv paper proposes Energy-Based Transformers (EBTs), a new class of Energy-Based Models that learn System 2-style thinking purely from unsupervised pretraining. EBTs assign an energy score to each input-candidate prediction pair, then reframe prediction as gradient descent-based energy minimization until convergence. Unlike existing inference-time compute methods, they need no verifiers or verifiable rewards and generalize across discrete (text) and continuous (visual) modalities. Across training, EBTs scale up to 35% faster than Transformer++ in terms of data, batch size, parameters, FLOPs, and depth.

Original post →

More from Research

Research channel →