Grok 4.5 Medium tops LaurenBench with 56.9%, ahead of Claude Sonnet 5 and GLM 5.2
elonmusk · x · 2026-07-29
A screenshot of LaurenBench shows Grok 4.5 Medium taking the top spot with 56.9%, ahead of Claude Sonnet 5 Medium at 55.6%, GLM 5.2 Medium at 55.2%, Claude Opus 5 Medium at 51.7%, and Kimi K3 Medium at 49.7%.
The benchmark is described as measuring real-world agent performance across conversation, tool use, memory, and safety. Farther down the list are GPT-5.6 Luna Medium at 43.7%, Gemini 3.5 Flash Lite Medium at 42.4%, Gemini 3.6 Flash Medium at 41.7%, and DeepSeek V4 Pro Medium at 41.3%.
More from Models
- Olmo 3 talk covers DPO, data problems, and how research reaches frontier models — natolambert · 2026-07-29
- Google DeepMind opens staff scientist role for Gemini long-context and memory — LucaAmb · 2026-07-29
- Claude Opus 5 climbs to No. 2 in Agent Arena, ahead of GPT-5.6 Sol — scaling01 · 2026-07-29
- Macaron-V1-Tall trends on Hugging Face as a text-generation model — mindlab-research · 2026-07-29
- Claude Opus 5 is listed at $5 in and $25 out per million tokens — arena · 2026-07-29
- Kimi's Open Model Pricing Sparks Debate: The $20M Monetization Reality — BenBajarin · 2026-07-29