Fixing Non-Autoregressive Generation with Zero Extra Compute: Perplexity Drops by 63%

kastnerkyle · x · 2026-08-02

Discrete flow models in non-autoregressive text generation often face an issue: they try to predict tokens even when the surrounding context is masked, forcing blind guesses and injecting noise into training.

This paper proposes a remarkably clean fix: reweight the loss and tweak the sampler so the model accounts for available local context before taking a loss penalty. Adding basically zero compute, this approach drops perplexity on OpenWebText by up to 63% and finally matches semi-autoregressive performance.

Original post →

More from Research

Research channel →