Flow Reasoning Models hit 99.5% on Sudoku-Extreme with 44x fewer inference FLOPs
burny_tech · x · 2026-09-06
A new paper on alphaXiv introduces Flow Reasoning Models (FRMs), a structured-reasoning framework that self-conditions a flow model on its own past outputs, turning one-shot denoising into iterative solution refinement — closer to repeated error correction than sequential generation.
To fix the exposure bias that breaks self-conditioning at greater recurrent depth, the authors propose Fixed-Point Forcing (FPF), which trains FRMs on states produced by their own inference dynamics while preserving the standard flow-matching objective.
Results: 99.5% solve rate on Sudoku-Extreme, 100.0% on Zebra, and 99.9% on Maze-Unique. On Sudoku-Extreme, FRMs beat evaluated masked-diffusion and specialized reasoning baselines on peak accuracy while matching the next-best method's 98.7% with 44x fewer inference FLOPs.
More from Research
- YC-backed MovingAtomsLab banned from DeepMind's Physics-IQ Verified benchmark for 3 months — HildeKuehne · 2026-09-06
- Is a schema-aware memory graph 'overfitting'? Dev asks for the cleanest leakage test — chaachans · 2026-09-06
- Burkov: 2026 is putting recurrence back into the Transformer it removed in 2017 — burkov · 2026-09-06
- MIT study: 83% of ChatGPT essay writers couldn't quote a single line they just wrote — victor_explore · 2026-09-06
- Carbon nanocone + fullerene check valve shows >10,000x rectification in MD sims — jwt0625 · 2026-09-06
- Formalize All Human Math in a Year? Bold AI Plan Gets Eric Weinstein's Backing — AccBalanced · 2026-09-06