Flow Reasoning Models hit 99.5% on extreme Sudoku via recurrent refinement
alec_helbling · x · 2026-09-03
Researchers including Duen Horng Chau and Mauro Martino publish 'Flow Reasoning Models: Turning Flows Into Efficient Recurrent Reasoners'. FRMs adapt continuous flows to discrete structured outputs and, by self-conditioning a flow model on its own past outputs, turn one-shot denoising into iterative solution refinement — enabling parallel, revisable interdependent decisions that autoregressive models (sequential, can't revise) and masked diffusion models (need careful decoding coordination) both struggle with.
The authors identify exposure bias that makes conventional self-conditioning unreliable at greater recurrent depth, and propose Fixed-Point Forcing (FPF): training on states produced by the model's own inference dynamics while keeping the standard flow-matching objective.
Results: solve rates of 99.5% / 100.0% / 99.9% on Sudoku-Extreme, Zebra, and Maze-Unique, with higher peak accuracy on Sudoku-Extreme than evaluated masked-diffusion baselines. Bonus: correct answers turn out to be fixed-point attractors of the self-conditioned flow's dynamics.
Related event: Flow Reasoning Models Solve Sudoku With 44x Less Compute(4 posts)→
More from Research
- Cohere Labs releases ATE dataset of ~700K MCP tools to reveal what agents do — Cohere_Labs · 2026-09-03
- Action Chunking Boosts Contrastive RL Even in Fully Online RL, Study Finds — ben_eysenbach · 2026-09-03
- Coding models are running out of data — PL researchers propose 'intent computing' as the fix — LingmingZhang · 2026-09-03
- davidad Backs Call to Ban Naive RLVR: 'Everything Should Be Model-Graded' — davidad · 2026-09-03
- Computerphile Deep Dive: How Watermarks Track AI-Generated Content — Computerphile · 2026-09-03
- TrafficLab 3D builds digital-twin traffic visualizations from CCTV footage and Google Maps — tom_doerr · 2026-09-03