LLaDA MoE v2 Matches Qwen3 with Only 65% Training Tokens

GSAI-ML · hf · 2026-08-05

This paper systematically characterizes the scaling behavior of Mixture-of-Experts (MoE) diffusion language models (dLLMs), identifying key differences from autoregressive (AR) models. Optimization requires faster batch size growth and quicker learning rate decay, while IsoFLOP analysis reveals a slight data-side tilt in compute allocation.

Guided by these findings, the authors trained LLaDA MoE v2 (30B-A3B) from scratch on 23.5T tokens. Using approximately 65% of Qwen3's pretraining tokens, it approaches Qwen3 on several knowledge and reasoning benchmarks. After SFT alone, it outperforms SDAR Chat on 7 out of 8 reasoning and coding benchmarks.

Original post →

More from Research

Research channel →