LLaDA 2.2’s block-diffusion MoE edits its own outputs during training and inference
omarsar0 · x · 2026-07-26
This follow-up explains the block-diffusion MoE design behind LLaDA 2.2.
- The diagram shows a three-stage pipeline: pretraining with Levenshtein editing, supervised fine-tuning on long-context and agent data, then RL with Levenshtein editing.
- It also shows the inference side: block routing selects a subset of experts per block to keep routing efficient.
- The core idea is to let the model edit its own spans instead of locking in early token mistakes.
More from Models
- Grok Voice is pitched as a 3x-faster alternative to typing for everyday work — Daniel_Farinax · 2026-07-28
- Arav Srinivas calls GLM underrated and says 700B feels close to Opus — AravSrinivas · 2026-07-28
- Kimi K3 reproduces RLVR findings without overclaiming, author says — infoxiao · 2026-07-28
- Inference.net pitches a gateway flow that mirrors prod traffic to Kimi K3 before switching — MatthewBerman · 2026-07-28
- Claude was unsubscribed as ChatGPT/Codex 5.6, Sol and Kimi 3 all struggled — sull · 2026-07-28
- GLM 5.5 is said to arrive in August with stronger long-horizon agent loops — bindureddy · 2026-07-28