GLM-5.2 nearly doubles Kimi K3 on a long-horizon browser benchmark
zainhas · x · 2026-07-24
A composite benchmark for long-horizon browser work shows GLM-5.2 well ahead of Kimi K3.
- The chart reports GLM-5.2 at 63% versus Kimi K3 at 31% on Composite-Bench.
- The benchmark targets long-horizon browser tasks with exact-match grading and pinned max reasoning effort.
- Other models in the chart include Claude Fable 5 (74%), Claude Opus 4.8 (73%), Gemini 3.5 Flash (27%), MiniMax-M3 (26%), GPT-5.6 Sol (18%), and Grok-4.5 (9%).
More from Models
- Grok 4.7 lands as OpenAI suspends $200 ChatGPT Pro plans — iannuttall · 2026-09-11
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11