New Long-Horizon Terminal Benchmark Released
minchoi · x · 2026-07-14
The newly released Long-Horizon Terminal-Bench is designed to test long-horizon terminal task capabilities, rather than short tasks that can be completed in a few minutes.
- The benchmark includes 46 real-world terminal tasks across 9 categories
- Evaluated 18 frontier models using the unified Terminus-2 harness
- A single task can last up to 90 minutes and involve roughly 120–320 steps
- Employs a hidden, replay-based verifier to prevent shortcuts and reward hacking
- Introduces continuous partial scoring to avoid compressing "progress made" into a simple 0/1
The author concludes that:
- The best model only achieves an average reward of 0.505
- 29/46 tasks remain unsolved by any model
- Long-horizon agents can currently "start working," but are still far from "getting the job done"
The post also mentions that Grok 4.5 currently tops the leaderboard, but overall, current models are still far from achieving truly reliable terminal execution.
Related event: Tencent Hunyuan Releases Long-Horizon Terminal Bench Exposing Model Limits(6 posts)→
More from Models
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Claude is no longer available for minors as Anthropic rolls out age assurance — Muhammad523 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11