Tencent Hunyuan Releases Long-Horizon Terminal Bench Exposing Model Limits

Tencent Hunyuan released Long-Horizon-Terminal-Bench to test the limits of agents in long-horizon terminal tasks. It is noteworthy because it shifts the evaluation focus from short tasks to real-world workflows requiring sustained, multi-step execution over extended periods.

Key Details

According to official posts, the benchmark includes 46 long-duration tasks across 9 scenarios. @_akhaliq noted that a core design feature is dense reward-based scoring, measuring task completion with finer granularity rather than just checking the final outcome. Multiple posts mention that a single task can take up to 90 minutes, and about 20 frontier models were evaluated in a unified environment.

Results and Signals

@iamfakhrealam summarized that the tests exposed clear shortcomings in current models' long-horizon terminal execution. Grok 4.5 ranked first, but the best model solved only 13 out of 46 tasks; 29 tasks remained unsolved by any model, and about 55% of runs scored below expectations. The overall signal is that large models still have a significant gap to bridge before they can reliably handle long-duration, multi-step terminal tasks requiring continuous operation and stability.

2026-07-13 ~ 2026-07-14 · 6 related posts