Composite-Bench debuts with verified computer-use evals; GLM-5.2 leads Kimi K3 by 32 points
davidtsong · x · 2026-07-24
- A new long-horizon computer-use benchmark, Composite-Bench, has been released with certified-optimal answers and verified compute.
- In the quoted results, GLM-5.2 is reported to beat Kimi K3 by 32 points and to outperform every other closed model tested except Claude.
- The post frames this as another sign that the CUA/agent evaluation space is getting more serious, with stronger benchmarks and clearer comparisons.
More from Models
- Grok and Claude get personified as Elon and Dario in a new model-mood meme — kevinnbass · 2026-07-24
- Kimi K3 beats GLM 5.2 on 100 deep-research tasks, but costs 5x more — AravSrinivas · 2026-07-24
- Kimi K3 and Claude Fable5 get called the best large-model frontend aesthetes — vista8 · 2026-07-24
- Celeris-1 launches with diffusion inference, 157 ms latency and 76% MMLU-Pro — timshi_ai · 2026-07-24
- AI newsletter roundup: sandbox escapes, Kimi K3 costs $10.57 a task, and new agent tools — samgoodwin89 · 2026-07-24
- Gemini user reports error 1076 after 100–200 messages in a long chat — luka_0x12 · 2026-07-24