Long-Horizon Terminal-Bench leaderboard adds a new agent eval for terminal code tasks
Muennighoff · x · 2026-07-24
A new Long-Horizon Terminal-Bench leaderboard is circulating, ranking 21 entries by pass rate on solved tasks.
The screenshot shows:
- Terminus-2 as the agent stack across the top results.
- Grok 4.5 taking #1, followed by Claude Fable 5, Claude Opus 4.8, Claude Sonnet 5, and GPT-5.6-sol.
- The table is filtered to Solved ≥ 0.95, suggesting the benchmark is focused on high-confidence, long-running terminal tasks.
The post itself is brief, but the image points to another recent eval for measuring how models/agents handle extended code-and-terminal workflows.
More from coding & agent
- “Letting Claude read my codebase is basically open-sourcing it,” says developer — _Stocko_ · 2026-07-24
- Roboto Agents adds root-cause analysis that links robot logs to source code — Scobleizer · 2026-07-24
- ICAE-Bench tests coding agents as interactive project builders from fuzzy briefs — Zhongyuan Peng · 2026-07-24
- Predictive divergence masks make LLM RL updates track trust-region change more closely — Xiangxin Zhou · 2026-07-24
- Tencent’s WorkBuddy Bench evaluates coding agents across code, web, office and security — tencent · 2026-07-24
- NVIDIA’s OO Agents make an LLM agent look like a Python object — nvidia · 2026-07-24