Representation-Space MMD Post-Training Boosts Diffusion LMs, More Parallel Decoding at 16B
yresearch · hf · 2026-10-06
yresearch introduces a post-training method for diffusion language models (DLMs) minimizing Maximum Mean Discrepancy (MMD) between generated and reference distributions in the feature space of a frozen pretrained DLM, retaining contextual token-level features for multiple observations per extractor pass.
The objective is optimized via policy gradients for discrete models and direct differentiation through generated latents for continuous ones—no full sampling trajectories or jointly trained auxiliary models needed.
Experiments show lower generative perplexity at comparable entropy on OpenWebText, better accuracy-compute trade-offs on GSM8K, and on 16B DMax-LLaDA2.0 (hybrid masked-uniform diffusion), increased decoding parallelism with similar or higher accuracy on math and code benchmarks.
More from Research
- Sander Dieleman on why continuous diffusion language models are making a comeback — LucaAmb · 2026-10-06
- Reading the source of 7 LLM eval tools uncovered 13 scoring bugs, 6 fixes merged — maverick_man1111 · 2026-10-06
- Terminal-Bench Pro: 400 tasks across 8 domains with zero contamination risk — thisguyknowsai · 2026-10-06
- ROME's training pipeline: 500B-token CPT, error-masked SFT, chunk-level RL — thisguyknowsai · 2026-10-06
- ROME's IPA assigns RL credit at chunk level, not token level, for tool-use agents — thisguyknowsai · 2026-10-06
- ROME team open-sourced full agent infra: ROLL, ROCK and iFlow CLI before training the model — thisguyknowsai · 2026-10-06