xAI Launches Grok 4.7: Same Price, Frontier Price-Performance With 46.3% on CursorBench
SpaceXAI · x · 2026-09-22
SpaceXAI officially launched Grok 4.7, its most capable coding and knowledge-work model, claiming twice the speed of comparable models at half the price — actual pricing matches Grok 4.6 ($2/M input, $6/M output tokens).
Model improvements
- New, larger base model with a longer RL run weighted toward multi-hour problems
- Better self-verification and long-context management
- Natively understands the Grok Bot harness
Key benchmarks (4.7 / 4.6 / GPT-5.6 Sol / Fable 5.1)
- CursorBench 4.0: 46.3% / 40.4% / 41.7% / 51.8%
- DeepSWE v1.1: 71.0% / 65.2% / 72.7% / 70.0%
- EEBench: 64.0% / 53.0% / 39.4% / 56.4%
- Terminal-Bench 4.0: 38.0% / 20.3% / 37.3% / 57.9%
- Harvey Legal: 19.6% / 15.8% / 2.5% / 6.7%
- HealthBench Professional: 56.7% / 48.5% / 60.5% / 62.1%
Frontier price-performance on CursorBench. Available now in Cursor, Grok Build, and the Grok API.
Related event: xAI Launches Grok 4.7, Its Strongest Coding Model Yet(26 posts)→
More from Models
- Azure OpenAI content filter blocks 'S&M' — the standard finance shorthand for Sales & Marketing — peterjliu · 2026-09-22
- IFM's K2-Horizon-36B-A4B Matches 20x-Larger Models on AA Index Using New MoVA Architecture — victormustar · 2026-09-22
- Grok 4.7 fails again: $1.59 run produces laughable output — teortaxesTex · 2026-09-22
- LLM scam detection benchmarked: fitted TF-IDF baseline beats Jev, DeepSeek and local Qwen — justinbiebar · 2026-09-22
- Goodfire Finds DNA Model Evo 2 Encodes the Tree of Life as a Curved Activation Manifold — burny_tech · 2026-09-22
- Internal eval puts Grok 4.7 at #3 across 22 knowledge-work tasks for under $5 — realsohamparekh · 2026-09-22