Long-Horizon Terminal-Bench leaderboard adds a new agent eval for terminal code tasks
Muennighoff · x · 2026-07-24
A new Long-Horizon Terminal-Bench leaderboard is circulating, ranking 21 entries by pass rate on solved tasks.
The screenshot shows:
- Terminus-2 as the agent stack across the top results.
- Grok 4.5 taking #1, followed by Claude Fable 5, Claude Opus 4.8, Claude Sonnet 5, and GPT-5.6-sol.
- The table is filtered to Solved ≥ 0.95, suggesting the benchmark is focused on high-confidence, long-running terminal tasks.
The post itself is brief, but the image points to another recent eval for measuring how models/agents handle extended code-and-terminal workflows.
More from coding & agent
- alphaXiv open-sources OpenResearch to run parallel research agents with any model — alphaXiv · 2026-09-11
- MathModelAgent gains traction: auto-solves math modeling and writes a submission-ready paper — jihe520 · 2026-09-11
- DeskcommCRM: open-source AI sales CRM with native agents and WhatsApp hits 1k stars — melgarafael · 2026-09-11
- hyperresearch: agent-driven knowledge base that turns web research into a searchable wiki — jordan-gibbs · 2026-09-11
- Forter's 13 lessons from its agent sprint: skip custom RAG, lean on mature enterprise search — bibryam · 2026-09-11
- Two real 'company brains' opened up live: Gorgias' in-house Cortex vs Slite — femke_plantinga · 2026-09-11