Long-Horizon Agent Progress on Terminal-Bench
tetsuoai · x · 2026-07-14
The post compares the performance of long-horizon agents on Terminal-Bench:
- Earlier papers concluded there was vast room for improvement on this benchmark: the best of 15 models completed only 7 out of 46 tasks, averaging around 2.
- The author now claims Grok 4.5 has achieved 13 tasks and Fable 5 has reached 12, indicating significant progress in long-horizon agent execution over the past two months.
- Cost and execution metrics per task include roughly 9.9 million tokens, 231 episodes, and 85 minutes of wall-clock time.
- The author analyzes that Grok 4.5's "4.2x output token efficiency" actually underestimates the gains, as the primary cost on such benchmarks stems from input replay. Completing tasks in fewer steps reduces the accumulated context, substantially lowering overall input costs.
- The post speculates that Grok 4.5's improvements might stem from Cursor-related training data, specifically a massive amount of real-world agentic editing trajectories.
Related event: Grok 4.5 Tops Long-Horizon Terminal-Bench(3 posts)→
More from coding & agent
- Building a Secure AI Agent Gateway: Self-Hosting OAuth for Multiple SaaS Apps — Defiant_Cod_2654 · 2026-07-22
- Rowboat launches as an open-source, local-first AI coworker with memory — ycombinator · 2026-07-22
- Reddit user chains Ideogram 4 and Krea2 to mimic bbox-based image positioning — v3lh0t05c0 · 2026-07-22
- Apollo Cuts AI Assistant Skill Dev Time by 85% with Deep Agents — LangChain · 2026-07-22
- Scoble says AI “loops” really means long-running multi-agent workspaces — Scobleizer · 2026-07-22
- Kimi Code opens a waitlist as Moonshot rolls out its coding product — Fabulous_Bonus_8981 · 2026-07-22