RUC & Ant Group's LLaDAMoEv2 Matches Qwen3 with 65% Pretraining Tokens

机器之心 · wechat · 2026-08-11

A joint team from RUC and Ant Group has systematically characterized the scaling laws for Mixture-of-Experts diffusion language models. Based on these laws, they trained a new 30B-A3B model, LLaDAMoEv2, from scratch.

The team summarized scaling laws specific to diffusion models across hyperparameters, compute allocation, and architecture design. Experiments show that LLaDAMoEv2, activating about 3B parameters, uses only 65% of the pretraining tokens of a similar-scale autoregressive model to approach Qwen3's performance on multiple benchmarks. In coding and reasoning tasks, the model surpassed existing leading diffusion models after just supervised fine-tuning, marking a shift from 'experience-driven' to 'law-guided' scaling for diffusion language models.

Original post →

More from Models

Models channel →