Grok 4.5 tops a long-horizon terminal benchmark, Elon Musk says

elonmusk · x · 2026-07-21

Elon Musk promotes Grok Build and cites a repost claiming Grok 4.5 ranked #1 on Long-Horizon Terminal-Bench by binary pass rate. - The quoted post says Grok 4.5 beat Claude Fable 5, Claude Opus 4.8, and GPT-5.6-sol under the strictest scoring metric. - The argument is that long-horizon terminal tasks stress full-workflow completion, recovery from mistakes, and sustained context over hundreds of steps. - The takeaway: Grok 4.5 appears particularly strong on complex agentic coding and automation work.

Original post →

More from coding & agent

coding & agent channel →