New Long-Horizon Terminal Benchmark Released
minchoi · x · 2026-07-14
The newly released Long-Horizon Terminal-Bench is designed to test long-horizon terminal task capabilities, rather than short tasks that can be completed in a few minutes.
- The benchmark includes 46 real-world terminal tasks across 9 categories
- Evaluated 18 frontier models using the unified Terminus-2 harness
- A single task can last up to 90 minutes and involve roughly 120–320 steps
- Employs a hidden, replay-based verifier to prevent shortcuts and reward hacking
- Introduces continuous partial scoring to avoid compressing "progress made" into a simple 0/1
The author concludes that:
- The best model only achieves an average reward of 0.505
- 29/46 tasks remain unsolved by any model
- Long-horizon agents can currently "start working," but are still far from "getting the job done"
The post also mentions that Grok 4.5 currently tops the leaderboard, but overall, current models are still far from achieving truly reliable terminal execution.
Related event: Tencent Hunyuan Releases Long-Horizon Terminal Bench Exposing Model Limits(6 posts)→
More from Models
- NVIDIA says Nemotron 3 Ultra scored 30/42 on the 2026 IMO problems — NVIDIAAI · 2026-07-22
- Gemma-4-26B-a4B reportedly beats Qwen3.6 and Qwen3.5 MoE fine-tunes — JLeonsarmiento · 2026-07-22
- OpenAI is reportedly briefing U.S. lawmakers on its next model family — kimmonismus · 2026-07-22
- Muse Spark 1.1 lands at 1495 on Text Arena with standout agentic-coding price performance — ycombinator · 2026-07-22
- Advanced AI Models Are Becoming Impossible to Plug and Play — emollick · 2026-07-22
- Google Gemini's AI Problem: No Leading Model for Core Workloads — bindureddy · 2026-07-22