LLaDA 2.2 aims at real agents with 1.64× throughput and stronger SWE-bench scores
omarsar0 · x · 2026-07-26
A diffusion LLM for agentic work, LLaDA 2.2 pairs block-parallel decoding with long-horizon planning and tool use.
- The model is presented as the first large-scale diffusion LLM built to act as a real agent, with self-correction over long multi-turn trajectories.
- Benchmark charts show strong results versus Ling-2.6-flash, including SWE-bench Verified 519.0 vs 303.2, SWE-bench Pro 485.3 vs 283.4, SWE-bench Multilingual 459.5 vs 200.6, and τ²-Bench 592.8 vs 334.9.
- The paper also claims 1.64× BF16 throughput over Ling-2.6-flash, with FP8 quantization adding another 18.6%.
- Technical details include native 128K context, MoE routing, Block Routing for predictable long-context cost, Levenshtein editing with KEEP/SUBSTITUTE/DELETE/INSERT, and an RL stage called L-EBPO to reduce trajectory-level error propagation.
More from Models
- OpenAI rated Astra 'Critical' for cyber capabilities — and admits it's harder to monitor — theguywhobuilds · 2026-09-11
- TestingCatalog's Daily AI Brief adds email editions, dishing Meta Muse and GPT-Live-1 rumors — testingcatalog · 2026-09-11
- ChatGPT monthly active users top 1.06 billion in August, fourth straight record month — FinanceYF5 · 2026-09-11
- PuzzleMask: Plain-Prose Attack Bypasses All 4 Tested LLM Gatekeepers at 100% — TechNadu · 2026-09-11
- OpenAI Codex may issue another usage reset this weekend, says Codex lead resets happen — umesh_ai · 2026-09-11
- OpenAI Reportedly Pointing Its Navier–Stokes Model at Riemann and P vs NP — 141_1337 · 2026-09-11