LLaDA 2.2 claims 703.82 TPS on BFCL-V4 and 519 TPS on SWE-bench Verified

FellMentKE · x · 2026-07-24

LLaDA 2.2 pushes diffusion LLMs toward agent use

The post says LLaDA 2.2 is an agent-oriented MoE diffusion LLM whose key improvement is self-correction during decoding via Levenshtein editing — deletion, insertion, and substitution. The claim is that for agents, the real bottleneck was error accumulation, not block-parallel decoding itself.

Reported performance

It is positioned as a high-throughput, low-latency engine for AI agents, with links to a GitHub repo, Hugging Face, and a technical report.

Related event: inclusionAI Launches LLaDA2.2-flash: 100B Diffusion Model for Agent Acceleration(9 posts)→

Original post →

More from Models

Models channel →