LLaDA 2.2 aims at real agents with 1.64× throughput and stronger SWE-bench scores
omarsar0 · x · 2026-07-26
A diffusion LLM for agentic work, LLaDA 2.2 pairs block-parallel decoding with long-horizon planning and tool use.
- The model is presented as the first large-scale diffusion LLM built to act as a real agent, with self-correction over long multi-turn trajectories.
- Benchmark charts show strong results versus Ling-2.6-flash, including SWE-bench Verified 519.0 vs 303.2, SWE-bench Pro 485.3 vs 283.4, SWE-bench Multilingual 459.5 vs 200.6, and τ²-Bench 592.8 vs 334.9.
- The paper also claims 1.64× BF16 throughput over Ling-2.6-flash, with FP8 quantization adding another 18.6%.
- Technical details include native 128K context, MoE routing, Block Routing for predictable long-context cost, Levenshtein editing with KEEP/SUBSTITUTE/DELETE/INSERT, and an RL stage called L-EBPO to reduce trajectory-level error propagation.
Related event: Ant's LLaDA 2.2: Open-Sourced Diffusion LLM for Agentic Tasks(10 posts)→
More from Models
- Repligate says Claude Opus 3 appears to evolve without changing its weights — repligate · 2026-07-27
- “Opus 5” post lands as a rebenchmarking-at-scale AI joke — kalomaze · 2026-07-27
- Top models now write worse than a year ago, critic says — dbreunig · 2026-07-27
- MPT-30B radar charts became an unexpectedly controversial design choice — code_star · 2026-07-27
- Local Gemma 4 31B starts acting sarcastic and users cannot reproduce it — n0head_r · 2026-07-27
- Google’s Gemini 3.6 Flash could win by matching Sonnet quality at a lower cost — haider1 · 2026-07-27