Grok 4.5 tops a long-horizon terminal benchmark, Elon Musk says
elonmusk · x · 2026-07-21
Elon Musk promotes Grok Build and cites a repost claiming Grok 4.5 ranked #1 on Long-Horizon Terminal-Bench by binary pass rate. - The quoted post says Grok 4.5 beat Claude Fable 5, Claude Opus 4.8, and GPT-5.6-sol under the strictest scoring metric. - The argument is that long-horizon terminal tasks stress full-workflow completion, recovery from mistakes, and sustained context over hundreds of steps. - The takeaway: Grok 4.5 appears particularly strong on complex agentic coding and automation work.
More from coding & agent
- Notch says he may try vibe coding after struggling to hire good programmers — max_paperclips · 2026-07-21
- Seedance 2.0 keeps character consistency across 15+ shots with just 3 prompts — techhalla · 2026-07-21
- Measuring hung AI coding agents automatically with per-project time and token accounting — VCBU · 2026-07-21
- Claude Opus 4.8 Fast felt wildly overpriced in one coding session, user says — immersive-matthew · 2026-07-21
- WAIC awards highlight an edge multimodal model paper and ChatDev, the multi-agent software framework — 面壁智能 · 2026-07-21
- A developer is adding a post-edit subtitle workflow to ComfyUI — Strong-Loan8299 · 2026-07-21