Amazon's ALoDLM: Token-Adaptive Looped Diffusion LMs Beat AR Baselines

amazon · hf · 2026-10-06

ALoDLM: Adaptively Looped Diffusion Language Models

Diffusion language models (DLMs) generate tokens in parallel for speed, but lag autoregressive (AR) models in quality. The paper attributes this to a computation-difficulty mismatch: some unknown tokens are easy to predict while others need far more compute, yet existing DLMs apply uniform depth to all positions at each denoising step.

ALoDLM replaces uniform computation with token-adaptive latent recurrence: at each step it iteratively refines latents and allocates compute by token difficulty—ready tokens are committed as discrete context, unresolved ones keep refining through extra recurrent passes. Token-wise computation schedules are formulated as latent variables, trained jointly via a conditional negative evidence lower bound (NELBO).

Results: trained at 1.7B and 8B scales, ALoDLM outperforms all evaluated DLMs and corresponding AR baselines in average score across eleven benchmarks, while retaining fast parallel decoding and a strong quality-efficiency trade-off under optimized inference engines.

Original post →

More from Models

Models channel →