Grok 4.7 Released: Terminal-Bench Score Nearly Doubles in Push for Multi-Hour Tasks
Scobleizer · x · 2026-09-22
SpaceXAI has released Grok 4.7, pitched as its strongest model for coding and knowledge work.
- New, larger base model with a longer RL run focused on multi-hour tasks
- Better self-verification, long-context handling, and native Grok Bot harness understanding
- Big benchmark jumps: CursorBench 4.0 from 40.4% to 46.3%, DeepSWE v1.1 from 65.2% to 71.0%, Terminal-Bench 4.0 from 20.3% to 38.0%
- Beats GPT-5.6 Sol and Fable 5.1 on several benchmarks
This is a third-party recap; see the official announcement for details.
Related event: xAI Launches Grok 4.7, Its Strongest Coding Model Yet(26 posts)→
More from Models
- Azure OpenAI content filter blocks 'S&M' — the standard finance shorthand for Sales & Marketing — peterjliu · 2026-09-22
- IFM's K2-Horizon-36B-A4B Matches 20x-Larger Models on AA Index Using New MoVA Architecture — victormustar · 2026-09-22
- Grok 4.7 fails again: $1.59 run produces laughable output — teortaxesTex · 2026-09-22
- LLM scam detection benchmarked: fitted TF-IDF baseline beats Jev, DeepSeek and local Qwen — justinbiebar · 2026-09-22
- Goodfire Finds DNA Model Evo 2 Encodes the Tree of Life as a Curved Activation Manifold — burny_tech · 2026-09-22
- Internal eval puts Grok 4.7 at #3 across 22 knowledge-work tasks for under $5 — realsohamparekh · 2026-09-22