Grok 4.5 Tops Long-Horizon Terminal-Bench
Grok 4.5 takes first place on the Long-Horizon Terminal-Bench, outperforming other frontier models like Claude Fable 5 and GPT-5.6 in complex agentic tasks.
2026-07-14 ~ 2026-07-15 · 3 related posts
- Long-Horizon Agent Progress on Terminal-Bench — tetsuoai · 2026-07-14
- GPT-5.6 Sol Leads in Agent Benchmarks — RajmaChawala · 2026-07-14
- Grok 4.5 Tops Terminal Agent Benchmark — XFreeze · 2026-07-15