LLaDA 2.2-flash posts 592.8 on τ²-Bench and 705.3 in fast mode
omarsar0 · x · 2026-07-26
More benchmark numbers for LLaDA 2.2-flash show its agentic and long-context performance.
- τ²-Bench: 592.80 vs 334.90 on Ling-2.6-flash.
- In fast mode, the score rises to 705.30.
- The attached table also compares throughput across SWE-bench, BFCL-v4, AIME 2026, LiveCodeBench, IFBench, KOR-Bench, GPQA-Diamond, and LongBench-v2, with the FP8-quantized variant usually faster still.
More from Models
- Grok Voice is pitched as a 3x-faster alternative to typing for everyday work — Daniel_Farinax · 2026-07-28
- Arav Srinivas calls GLM underrated and says 700B feels close to Opus — AravSrinivas · 2026-07-28
- Kimi K3 reproduces RLVR findings without overclaiming, author says — infoxiao · 2026-07-28
- Inference.net pitches a gateway flow that mirrors prod traffic to Kimi K3 before switching — MatthewBerman · 2026-07-28
- Claude was unsubscribed as ChatGPT/Codex 5.6, Sol and Kimi 3 all struggled — sull · 2026-07-28
- GLM 5.5 is said to arrive in August with stronger long-horizon agent loops — bindureddy · 2026-07-28