New Long-Horizon Terminal Benchmark Ranks Minimax, Kimi, and GLM Top 3

pcuenq · x · 2026-08-19

A new benchmark, Long-Horizon Terminal-Bench (LHTB), is released on the Hugging Face Hub to evaluate if LLM agents can maintain context over 300+ steps in a terminal. It features 46 tasks, contamination resistance, and hidden verifiers checking real system state. The current leaderboard is led by MiniMax AI's Minimax M3, followed by Moonshot's Kimi k2.7 Code and Zhipu's GLM 5.2.

Original post →

More from Research

Research channel →