Grok 4.5 tops a long-horizon terminal benchmark, Elon Musk says
elonmusk · x · 2026-07-21
Elon Musk promotes Grok Build and cites a repost claiming Grok 4.5 ranked #1 on Long-Horizon Terminal-Bench by binary pass rate.
- The quoted post says Grok 4.5 beat Claude Fable 5, Claude Opus 4.8, and GPT-5.6-sol under the strictest scoring metric.
- The argument is that long-horizon terminal tasks stress full-workflow completion, recovery from mistakes, and sustained context over hundreds of steps.
- The takeaway: Grok 4.5 appears particularly strong on complex agentic coding and automation work.
More from coding & agent
- Dev builds talk on guardrails workflow for shipping AI-written code without reading it — TejasKumar_ · 2026-09-11
- banteg: Codex auto-review has regressed, blocking steps needed to complete authorized tasks — banteg · 2026-09-11
- A doc-anchored agent workflow: you write, the agent only critiques and finds disagreements — lucasmeijer · 2026-09-11
- SymKit MCP: 44 tools for AI agents to verify symbolic derivations — Foreign-Specific-604 · 2026-09-11
- GitHub Copilot team routes user bug reports to an AI agent via Slack — marlene_zw · 2026-09-11
- Scanning 23 agent sessions, a dev found 3 silent failure modes in memory systems — No_Advertising2536 · 2026-09-11