Grok triples Terminal-Bench 4.0 score in two months, overtakes GPT-5.6 Sol

ns123abc · x · 2026-09-22

Grok reportedly tripled its Terminal-Bench 4.0 score in just two months, jumping from 12.4% to 38.0% and overtaking GPT-5.6 Sol on the terminal/agent coding benchmark. If accurate, it marks a rapid leap for xAI on agentic coding tasks.

Related event: Grok 4.7 Tops Terminal-Bench 4.0, Tripling Score in Two Months(2 posts)→

Original post →

More from Models

Models channel →