Flow Reasoning Models hit 99.5% on extreme Sudoku via recurrent refinement

alec_helbling · x · 2026-09-03

Researchers including Duen Horng Chau and Mauro Martino publish 'Flow Reasoning Models: Turning Flows Into Efficient Recurrent Reasoners'. FRMs adapt continuous flows to discrete structured outputs and, by self-conditioning a flow model on its own past outputs, turn one-shot denoising into iterative solution refinement — enabling parallel, revisable interdependent decisions that autoregressive models (sequential, can't revise) and masked diffusion models (need careful decoding coordination) both struggle with.

The authors identify exposure bias that makes conventional self-conditioning unreliable at greater recurrent depth, and propose Fixed-Point Forcing (FPF): training on states produced by the model's own inference dynamics while keeping the standard flow-matching objective.

Results: solve rates of 99.5% / 100.0% / 99.9% on Sudoku-Extreme, Zebra, and Maze-Unique, with higher peak accuracy on Sudoku-Extreme than evaluated masked-diffusion baselines. Bonus: correct answers turn out to be fixed-point attractors of the self-conditioned flow's dynamics.

Related event: Flow Reasoning Models Solve Sudoku With 44x Less Compute(4 posts)→

Original post →

More from Research

Research channel →