LLaDA 2.2 claims 1.64× throughput and native 128K context
omarsar0 · x · 2026-07-26
LLaDA 2.2’s practical selling points are speed and long-context stability.
- The model averages 1.64× BF16 throughput over Ling-2.6-flash.
- FP8 quantization adds another 18.6%.
- It supports native 128K context and was trained with a progressive schedule from 8K to 128K.
- It uses MoE routing so only a subset of parameters runs per token, and Block Routing further reduces routing overhead to keep long-context inference cost predictable.
More from Models
- Grok Voice is pitched as a 3x-faster alternative to typing for everyday work — Daniel_Farinax · 2026-07-28
- Arav Srinivas calls GLM underrated and says 700B feels close to Opus — AravSrinivas · 2026-07-28
- Kimi K3 reproduces RLVR findings without overclaiming, author says — infoxiao · 2026-07-28
- Inference.net pitches a gateway flow that mirrors prod traffic to Kimi K3 before switching — MatthewBerman · 2026-07-28
- Claude was unsubscribed as ChatGPT/Codex 5.6, Sol and Kimi 3 all struggled — sull · 2026-07-28
- GLM 5.5 is said to arrive in August with stronger long-horizon agent loops — bindureddy · 2026-07-28