Transformers proven computationally universal, but standard optimizers can't find the optimal state

Promptmethus · x · 2026-10-09

A new study applies Kolmogorov complexity—the mathematical limit of perfect compression and algorithmic generalization—to mathematically prove that Transformer encoders are computationally universal: there exists a state where the model is perfectly compressed and generalizes flawlessly.

The catch: when researchers built an objective to force the model toward this optimal state, the Transformer could hold it, but standard optimizers—the same ones used to train ChatGPT, Claude, and Gemini—simply cannot find it.

Takeaway: the architecture is theoretically flawless; the bottleneck is the optimization process, reshaping how we should view every AI lab's training methods.

Original post →

More from Research

Research channel →