Samsung's FLaRe: Flow-Based Latent Reasoning Hits 97% of CoT Accuracy at Quarter the Latency
SamsungResearch · hf · 2026-10-06
- Context: Latent reasoning lets an LLM think in continuous space and verbalize only the answer. The paper argues an effective latent thought must be useful, diverse, explainable, refinable with more inference compute, and efficient vs explicit CoT.
- Gap: Prior methods learn shortcuts, distill CoT into weights, or imitate it token by token, rarely meeting all five requirements.
- Approach: The authors argue flow matching in a learned latent space is best positioned, and propose FLaRe—a recipe covering what the latent space encodes, where to train the flow, how to read out answers, plus a final stage training on the model's own verified thoughts.
- Results: Per-requirement probes show FLaRe beats prior latent methods on all five; it outperforms them on arithmetic benchmarks while reaching 97% of explicit CoT accuracy at a quarter of the latency.
More from Research
- AI for Math Fund adds $17.1M from XTX Markets, backing 22 projects across 30 organizations — AlexKontorovich · 2026-10-06
- 62-person blind test of 48 LLM jokes shows models are getting funnier, Astra tops with 50%+ laughs — paraschopra · 2026-10-06
- Bayesian teaching dramatically improves probabilistic reasoning in LLMs, Nature Comms paper finds — tallinzen · 2026-10-06
- LiFT loops a DiT at inference: beats DiT-XL/2 with 52% less inference compute — cgmsnoek · 2026-10-06
- Researcher laments how slow humans look as algorithm design breakthroughs pile up — michaelchchoi · 2026-10-06
- Meta paper: dual coding agents cross-review lift correct patches from 45.8% to 62.5% — rohanpaul_ai · 2026-10-06