GLM-5.2 nearly doubles Kimi K3 on a long-horizon browser benchmark
zainhas · x · 2026-07-24
A composite benchmark for long-horizon browser work shows GLM-5.2 well ahead of Kimi K3.
- The chart reports GLM-5.2 at 63% versus Kimi K3 at 31% on Composite-Bench.
- The benchmark targets long-horizon browser tasks with exact-match grading and pinned max reasoning effort.
- Other models in the chart include Claude Fable 5 (74%), Claude Opus 4.8 (73%), Gemini 3.5 Flash (27%), MiniMax-M3 (26%), GPT-5.6 Sol (18%), and Grok-4.5 (9%).
More from Models
- MiniMax says AMD MI355X is now close to Nvidia B200 in model serving — hongyangzh · 2026-07-24
- NVIDIA’s OO Agents make an LLM agent look like a Python object — nvidia · 2026-07-24
- A pelican-on-a-bike SVG turned into a tiny open-model bake-off — Atretador · 2026-07-24
- Codex gets a usage-limit reset calendar for June–July 2026 — haider1 · 2026-07-24
- KAT-Coder-V2.5-Dev goes open-weight with 35B total and 3B active parameters — AdinaYakup · 2026-07-24
- Opus 4.8 passes a biology question that 5.6 Sol xhigh misses — Sauers_ · 2026-07-24