New Long-Horizon Terminal Benchmark Ranks Minimax, Kimi, and GLM Top 3
pcuenq · x · 2026-08-19
A new benchmark, Long-Horizon Terminal-Bench (LHTB), is released on the Hugging Face Hub to evaluate if LLM agents can maintain context over 300+ steps in a terminal. It features 46 tasks, contamination resistance, and hidden verifiers checking real system state. The current leaderboard is led by MiniMax AI's Minimax M3, followed by Moonshot's Kimi k2.7 Code and Zhipu's GLM 5.2.
More from Research
- Physics of Agents: Predicting Collective Behavior with Statistical Physics — james_y_zou · 2026-08-20
- The Human-or-Machine Issue: Turing-Inspired Reflections — ArtificialOther · 2026-08-20
- Meta Research Challenges Chinchilla Scaling Laws on Data-Compute Interactions — burkov · 2026-08-20
- 14,472 AI citations analyzed: business websites still win 60% of local search citations — gaganghotra_ · 2026-08-20
- LEGO-RL: harness-native reinforcement learning for coding agents — Lego-X · 2026-08-20
- Fourier Neural Operators predict quantum dynamics 10^7x faster than CUDA-Q — AnimaAnandkumar · 2026-08-20