Representation-Space MMD Post-Training Boosts Diffusion LMs, More Parallel Decoding at 16B

yresearch · hf · 2026-10-06

yresearch introduces a post-training method for diffusion language models (DLMs) minimizing Maximum Mean Discrepancy (MMD) between generated and reference distributions in the feature space of a frozen pretrained DLM, retaining contextual token-level features for multiple observations per extractor pass.

The objective is optimized via policy gradients for discrete models and direct differentiation through generated latents for continuous ones—no full sampling trajectories or jointly trained auxiliary models needed.

Experiments show lower generative perplexity at comparable entropy on OpenWebText, better accuracy-compute trade-offs on GSM8K, and on 16B DMax-LLaDA2.0 (hybrid masked-uniform diffusion), increased decoding parallelism with similar or higher accuracy on math and code benchmarks.

Original post →

More from Research

Research channel →