Flow Reasoning Models hit 99.5% on Sudoku-Extreme with 44x fewer inference FLOPs

burny_tech · x · 2026-09-06

A new paper on alphaXiv introduces Flow Reasoning Models (FRMs), a structured-reasoning framework that self-conditions a flow model on its own past outputs, turning one-shot denoising into iterative solution refinement — closer to repeated error correction than sequential generation.

To fix the exposure bias that breaks self-conditioning at greater recurrent depth, the authors propose Fixed-Point Forcing (FPF), which trains FRMs on states produced by their own inference dynamics while preserving the standard flow-matching objective.

Results: 99.5% solve rate on Sudoku-Extreme, 100.0% on Zebra, and 99.9% on Maze-Unique. On Sudoku-Extreme, FRMs beat evaluated masked-diffusion and specialized reasoning baselines on peak accuracy while matching the next-best method's 98.7% with 44x fewer inference FLOPs.

Related event: Flow Reasoning Models Rewrite Solutions Iteratively, Cutting Inference FLOPs 44x(3 posts)→

Original post →

More from Research

Research channel →