CMU and Oxford paper: refining hidden state beats longer chain-of-thought for reasoning

rohanpaul_ai · x · 2026-09-18

A new Carnegie Mellon + Oxford paper shows models can 'think longer' by repeatedly refining their hidden state instead of generating longer chains of thought—test-time compute doesn't have to mean more tokens. Looped flows fix the instability of long recurrent loops by training each update on a small denoising task while keeping the hidden state useful for the next update, letting models keep improving the same internal representation at inference time and materially raising accuracy.

Related event: CMU and Oxford: Looping Hidden States Can Replace Long Chain-of-Thought(2 posts)→

Original post →

More from Research

Research channel →