Grok 4.7 launches: 46.3% on CursorBench, Terminal-Bench nearly doubles, same $2/$6 pricing
tetsuoai · x · 2026-09-22
Grok 4.7 Launch Highlights
SpaceXAI (xAI) released Grok 4.7, its most capable model for coding and knowledge work, claiming 2x the speed at half the price of comparable models.
Model improvements
- New, larger base model with a longer RL run on a harder task mix weighted toward multi-hour problems
- Better self-verification and long-context management
- Natively trained to understand the Grok Bot harness for conversational and knowledge work
Benchmarks (vs GPT-5.6 Sol Max, Fable 5.1 Max)
- CursorBench 4.0: 46.3% (4.6: 40.4%, GPT-5.6 Sol: 41.7%)
- Terminal-Bench 4.0: 38.0%, nearly doubling from 20.3%
- EEBench: 64.0%, a wide lead
- Harvey Legal Agent Benchmark: 19.6%, far ahead of rivals
- HealthBench Professional: 56.7%, still behind Fable 5.1's 62.1%
Pricing & availability
- Same price as 4.6: $2/M input, $6/M output tokens; Fast variant doubles output speed at 2x price
- Available today in Cursor, Grok Build, and the Grok API
Related event: xAI launches Grok 4.7: same price, faster speed, big benchmark gains(16 posts)→
More from Models
- Grok 4.7 matches Opus 5 Max on coding at less than half the cost — XFreeze · 2026-09-22
- Grok 4.7 comparison clip shows major gains in 3D modeling and game physics over 4.6 — belce_dogru · 2026-09-22
- Hands-on: TypeSafe's tiny Jev classifier goes head-to-head with Claude Haiku/Sonnet/Opus on text understanding — vesko_st · 2026-09-22
- Hand-built anti-memorization benchmark: tiny Jev classifier matches Sonnet at ~150x lower cost — vesko_st · 2026-09-22
- HLE leaderboard: Grok 4.7 at #21 while Gemini 3.8 and Muse 1.3 lead by a margin — himanshustwts · 2026-09-22
- Grok 4.7's Terminal-Bench 4.0 coding score is 'horrendous', falling far behind OpenAI and Anthropic — daniel_mac8 · 2026-09-22