Long-Horizon Terminal-Bench leaderboard adds a new agent eval for terminal code tasks

Muennighoff · x · 2026-07-24

A new Long-Horizon Terminal-Bench leaderboard is circulating, ranking 21 entries by pass rate on solved tasks.

The screenshot shows:

The post itself is brief, but the image points to another recent eval for measuring how models/agents handle extended code-and-terminal workflows.

Original post →

More from coding & agent

coding & agent channel →