Grok 4.5 Medium tops LaurenBench with 56.9%, ahead of Claude Sonnet 5 and GLM 5.2

elonmusk · x · 2026-07-29

A screenshot of LaurenBench shows Grok 4.5 Medium taking the top spot with 56.9%, ahead of Claude Sonnet 5 Medium at 55.6%, GLM 5.2 Medium at 55.2%, Claude Opus 5 Medium at 51.7%, and Kimi K3 Medium at 49.7%.

The benchmark is described as measuring real-world agent performance across conversation, tool use, memory, and safety. Farther down the list are GPT-5.6 Luna Medium at 43.7%, Gemini 3.5 Flash Lite Medium at 42.4%, Gemini 3.6 Flash Medium at 41.7%, and DeepSeek V4 Pro Medium at 41.3%.

Original post →

More from Models

Models channel →