Grok 4.7's Terminal-Bench 4.0 coding score is 'horrendous', falling far behind OpenAI and Anthropic

daniel_mac8 · x · 2026-09-22

Grok 4.7 is out, but its Terminal-Bench 4.0 score is 'horrendous', according to the author — evidence of how far ahead OpenAI and Anthropic are on coding models. He still enjoys using Grok in Grok Bot for chat, just not for terminal coding tasks.

Original post →

More from Models

Models channel →