Grok 4.7 Tops Terminal-Bench 4.0, Tripling Score in Two Months

xAI's Grok 4.7 model card shows its Terminal-Bench 4.0 score jumping to 38.0% from 20.3%, surpassing GPT-5.6 Sol, though Claude Fable 5.1 still leads at 57.9%.

2026-09-22 ~ 2026-09-22 · 2 related posts