LLaDA 2.2-flash shows broad benchmark gains over Ling-2.6-flash
omarsar0 · x · 2026-07-26
The benchmark chart highlights LLaDA 2.2-flash’s gains across agentic and coding tasks.
- SWE-bench Verified: 49.28 vs 31.88 in the highlighted comparison.
- The same graphic also shows strong lead margins on SWE-bench Pro, SWE-bench Multilingual, τ²-Bench, BFCL-v4, MCP-Atlas, PinchBench, Claw-Eval, and related tasks.
More from Models
- Grok Voice is pitched as a 3x-faster alternative to typing for everyday work — Daniel_Farinax · 2026-07-28
- Arav Srinivas calls GLM underrated and says 700B feels close to Opus — AravSrinivas · 2026-07-28
- Kimi K3 reproduces RLVR findings without overclaiming, author says — infoxiao · 2026-07-28
- Inference.net pitches a gateway flow that mirrors prod traffic to Kimi K3 before switching — MatthewBerman · 2026-07-28
- Claude was unsubscribed as ChatGPT/Codex 5.6, Sol and Kimi 3 all struggled — sull · 2026-07-28
- GLM 5.5 is said to arrive in August with stronger long-horizon agent loops — bindureddy · 2026-07-28