Terminal-bench: Kimi k3 Tops Leaderboard, DeepSeek V4 Pro Second
SergioPaniego · x · 2026-08-24
Terminal-bench, featuring 89 high-quality tasks (e.g., fixing OCaml GC bugs, recovering SQLite DBs), is regarded as a key benchmark driving the field towards efficient, local agents.
Current Leaderboard:
- 🥇 Kimi k3
- 🥈 DeepSeek V4 Pro
- 🥉 Alibaba Qwen Max
- Honorable Mention: Qwen 3.8 27B, noted as the best model under 128B parameters.
More from Models
- Minimax H3 quality issues on RTX 4090 but not RTX 5090 — FoxTrotte · 2026-08-24
- Ox Alpha excels at Lean formalization — aiamblichus · 2026-08-24
- Stealth Startup's Continual Learning Model: Infinite Context & 15-Year Memory Tested — khademinori · 2026-08-24
- User praises Qwen 3.8 for versatility and hardware compatibility — uwukko · 2026-08-24
- Laguna XS 2.1: Top coding pick for low-VRAM GPUs — needthosepylons · 2026-08-24
- OX-Alpha outperforms DeepSeek-V4 in coding benchmarks: Test — karminski3 · 2026-08-24