FLM Does One-Step Language Modeling via Continuous Denoising, 8.3x Faster, NeurIPS 2026

ArashVahdat · x · 2026-09-25

A KAIST/CMU paper accepted at NeurIPS 2026 challenges the assumption that discrete diffusion is the right path for fast text generation. FLM uses flow-based continuous denoising on one-hot tokens, outperforming SOTA discrete diffusion models in many-step regimes; its distilled FMLM achieves few-step parallel generation that matches 8-step baseline quality in a single step, with an 8.3x speedup on LM1B and OpenWebText.

Original post →

More from Research

Research channel →