Grok 4.5 Tops Terminal Agent Benchmark
XFreeze · x · 2026-07-15
Grok 4.5 secured first place on the Long-Horizon Terminal-Bench, outperforming Claude Fable 5, Claude Opus 4.8, and GPT-5.6-sol.
This benchmark tests whether an AI agent can continuously advance tasks without losing context during up to 90 minutes of terminal operations involving hundreds of dependent steps. The author notes that across 46 high-difficulty tasks and 18 frontier models, Grok 4.5 achieved the highest average reward of 0.505.
The post emphasizes that this demonstrates Grok 4.5's superior performance in complex agentic workflows, real-world coding automation, and difficult engineering tasks.
Related event: Grok 4.5 Tops Long-Horizon Terminal-Bench(3 posts)→
More from coding & agent
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- First-ever Three.js Conference lands in Paris, with a panel on AI-shortened design workflows — OdinLovis · 2026-09-11
- Data engineering, not agent frameworks, is the real bottleneck for enterprise AI agents — dhruv2038 · 2026-09-11
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- Investment Analyst Asks How to Build a Claude-Based Diligence Agent Stack — Careless_Tie2286 · 2026-09-11