Grok 4.7 model card: Terminal-Bench 4.0 score jumps from 20.3% to 38.0%, still trails Claude Fable 5.1
ns123abc · x · 2026-09-22
xAI's Grok 4.7 model card is out, and the Terminal-Bench 4.0 leaderboard shows a big shakeup:
- Claude Fable 5.1 dominates at 57.9%, well clear of the field
- Grok 4.7 jumped from 20.3% to 38.0%, overtaking GPT-5.6 Sol
A substantial gain for Grok, though it still trails the leader by roughly 20 percentage points.
More from Models
- xAI's Grok hype cycle repeats: 4.7 delayed, no frontier model beats ChatGPT or Claude — flowersslop · 2026-09-22
- xAI fixes SDK bug dropping reasoning content, significantly boosting Grok 4.7 — ns123abc · 2026-09-22
- TTS leaderboard: xAI hits 87.6% pronunciation accuracy, Kokoro 82M fastest at 242 chars/s — ArtificialAnlys · 2026-09-22
- xAI's Post-Launch SDK Update Pushes It to #10 on the Vals Index — teortaxesTex · 2026-09-22
- Grok 4.7 looks pricier than before, now costlier than Astra on Artificial Analysis — steipete · 2026-09-22
- Which LLM is most encyclopedic on 8GB VRAM + 64GB RAM? — Mangleus · 2026-09-22