Long-Horizon Terminal-Bench: LLMs Still Struggle

iamfakhrealam · x · 2026-07-13

Long-Horizon Terminal-Bench tested 20 frontier models across 46 terminal tasks, with a 90-minute limit per task. Results show that even the best-performing model solved only 13 tasks, 29 tasks were unsolved by any model, and about 55% of runs scored below expectations. This indicates that LLMs still face significant challenges when handling long-horizon, real-world tasks.

Related event: Tencent Hunyuan Releases Long-Horizon Terminal Bench Exposing Model Limits(6 posts)→

Original post →

More from Models

Models channel →