Terminal Bench 4.0: one of the few evals whose ranking actually reflects capability
JJitsev · x · 2026-09-10
JJitsev says Terminal Bench 4.0 and Terminal Bench Science 0.1 are among the very few evals where he trusts the ranking reflects actual model capabilities, with still a healthy gap to saturation — though that may change with the next model generation within 6–12 months. He calls for community contributions to keep the open-source work high quality, as a new SOTA just landed on the benchmark.
More from Models
- OpenAI launches Sketch input for ChatGPT: draw instead of describe — floguo · 2026-09-10
- GPT 6 to 7 before GTA 6's Nov 19 launch would be OpenAI's fastest whole-number jump ever — ChrisGPT · 2026-09-10
- Why LLMs miscount the r's in strawberry: counting demands step-by-step enumeration — ctjlewis · 2026-09-10
- Sarvam Ships Realtime Streaming STT API With Mid-Call Reconfig and Millisecond VAD Tuning — itsOmSarraf_ · 2026-09-10
- Testing Astra's self-driven creativity with tree-search prompting: better variety, still lackluster — creatoroff · 2026-09-10
- Dev builds fun AI game with DeepSeek 4.1 Flash, praising its 'gamer temperament' — teortaxesTex · 2026-09-10