LLaDA MoE v2 Matches Qwen3 with Only 65% Training Tokens
GSAI-ML · hf · 2026-08-05
This paper systematically characterizes the scaling behavior of Mixture-of-Experts (MoE) diffusion language models (dLLMs), identifying key differences from autoregressive (AR) models. Optimization requires faster batch size growth and quicker learning rate decay, while IsoFLOP analysis reveals a slight data-side tilt in compute allocation.
Guided by these findings, the authors trained LLaDA MoE v2 (30B-A3B) from scratch on 23.5T tokens. Using approximately 65% of Qwen3's pretraining tokens, it approaches Qwen3 on several knowledge and reasoning benchmarks. After SFT alone, it outperforms SDAR Chat on 7 out of 8 reasoning and coding benchmarks.
More from Research
- Peking University Introduces ContinualSkillBench: Evaluating Continual Skill Evolution in LLM Agents — PekingUniversity · 2026-08-05
- AI Singapore Compresses LLM Training to 2 Days, Adds Five Low-Resource SEA Languages — davlanade · 2026-08-05
- NeurIPS Peer Review in Decline: ChatGPT Responses and Hallucinated Citations Plague Submissions — Pseudomanifold · 2026-08-05
- HSWQ NVFP4 Quantization for SDXL Significantly Outperforms Native Approach — Zestyclose_Bake3680 · 2026-08-05
- Dev Trains 203M Portuguese LLM from Scratch in 2.5 Hours on Single H100 — War_Enterprise · 2026-08-05
- NeurIPS Peer Review Debate: Should Reviewers Reveal Score Flexibility? — miniapeur · 2026-08-05