Grok 4.7's Terminal-Bench 4.0 coding score is 'horrendous', falling far behind OpenAI and Anthropic
daniel_mac8 · x · 2026-09-22
Grok 4.7 is out, but its Terminal-Bench 4.0 score is 'horrendous', according to the author — evidence of how far ahead OpenAI and Anthropic are on coding models. He still enjoys using Grok in Grok Bot for chat, just not for terminal coding tasks.
More from Models
- Alignment backfires: model strips all faces from a deepfake detection dataset mid-task — generativist · 2026-09-22
- Grok 4.7 falls to #24 on Vals Index, down 5 points from Grok 4.6 — scaling01 · 2026-09-22
- Game Theory of Model Launch Dates: Launching Early Admits Your Model Is Weaker — cocktailpeanut · 2026-09-22
- Liquid AI's LFM2.5 tops mobile benchmarks: 2.32GB memory, 8s latency on iPhone 17 Pro — maximelabonne · 2026-09-22
- Jev reportedly does tensor logic under the hood: differentiable IF args, no wasted gen tokens — StewartalsopIII · 2026-09-22
- Multilingual Model Laya Trending on Hugging Face — convaiinnovations · 2026-09-22