RUC & Ant Group's LLaDAMoEv2 Matches Qwen3 with 65% Pretraining Tokens
机器之心 · wechat · 2026-08-11
A joint team from RUC and Ant Group has systematically characterized the scaling laws for Mixture-of-Experts diffusion language models. Based on these laws, they trained a new 30B-A3B model, LLaDAMoEv2, from scratch.
The team summarized scaling laws specific to diffusion models across hyperparameters, compute allocation, and architecture design. Experiments show that LLaDAMoEv2, activating about 3B parameters, uses only 65% of the pretraining tokens of a similar-scale autoregressive model to approach Qwen3's performance on multiple benchmarks. In coding and reasoning tasks, the model surpassed existing leading diffusion models after just supervised fine-tuning, marking a shift from 'experience-driven' to 'law-guided' scaling for diffusion language models.
More from Models
- UCL Researchers Use RL to Stop LLMs from Lying About Their Hidden Reasoning — alex_verem · 2026-08-11
- Falcon-Perception: A 0.6B Model for Generating Labels for Object Detection and Segmentation — vanstriendaniel · 2026-08-11
- Claude Outputs Now Include Text Watermarks, Sparking Removal Discussions — Franck_Dernoncourt · 2026-08-11
- Predictions: Major Update by Late Sept, Next Paper to Focus on Agents — teortaxesTex · 2026-08-11
- Anthropic to embed invisible watermarks in all Claude text outputs globally — The Decoder · 2026-08-11
- Rumor: DeepSeek Holding onto V4 Pro GA Release — dejavucoder · 2026-08-11