Grok triples Terminal-Bench 4.0 score in two months, overtakes GPT-5.6 Sol
ns123abc · x · 2026-09-22
Grok reportedly tripled its Terminal-Bench 4.0 score in just two months, jumping from 12.4% to 38.0% and overtaking GPT-5.6 Sol on the terminal/agent coding benchmark. If accurate, it marks a rapid leap for xAI on agentic coding tasks.
Related event: Grok 4.7 Tops Terminal-Bench 4.0, Tripling Score in Two Months(2 posts)→
More from Models
- Grok ships three frontier models in 9 weeks: 4.5, 4.6 and 4.7 back-to-back — XFreeze · 2026-09-22
- OpenAI reportedly rushing Codex Bot to rival Grok Bot, warns on recursive self-improvement — dotey · 2026-09-22
- Engineer autonomously trains a Jev-competitive model with an agent swarm for $3.1k in 20 hours — denisyarats · 2026-09-22
- Challenge: Track Your Daily Token Usage to Prove AI Companies Are Throttling Limits — tomchapin · 2026-09-22
- Day 1 with Grok 4.7: strict system-prompt adherence and visible gains over 4.5 in real coding work — elonmusk · 2026-09-22
- MoVA adds sparse value experts to attention for more capacity at no extra KV-cache cost — rupspace · 2026-09-22