Unverified: Grok 4.7 Leak Claims Same Price, Terminal-Bench Jump From 20.3% to 38.0%
mark_k · x · 2026-09-22
An unverified post claims xAI's Grok 4.7 is out: a larger base with longer RL training on multi-hour agent tasks, at Grok 4.6's price ($2/M input, $6/M output; 2x-speed variant at 2x price). Claimed benchmark gains: CursorBench 40.4%→46.3%, Terminal-Bench 20.3%→38.0%, EEBench 53.0%→64.0%, Harvey 15.8%→19.6%, GDPval 1605→1695 Elo. It allegedly beats GPT-5.6 Sol on CursorBench/Terminal-Bench and leads Sol and Fable 5.1 on EEBench/Harvey, with a rebuilt safeguard system. Available in Cursor, Grok Build and the API. Note: the @SpaceXAI attribution and competitor names warrant skepticism.
Related event: Grok 4.7 rolls out at same price with big gains on long-horizon tasks(5 posts)→
More from Models
- Azure OpenAI content filter blocks 'S&M' — the standard finance shorthand for Sales & Marketing — peterjliu · 2026-09-22
- IFM's K2-Horizon-36B-A4B Matches 20x-Larger Models on AA Index Using New MoVA Architecture — victormustar · 2026-09-22
- Grok 4.7 fails again: $1.59 run produces laughable output — teortaxesTex · 2026-09-22
- LLM scam detection benchmarked: fitted TF-IDF baseline beats Jev, DeepSeek and local Qwen — justinbiebar · 2026-09-22
- Goodfire Finds DNA Model Evo 2 Encodes the Tree of Life as a Curved Activation Manifold — burny_tech · 2026-09-22
- Internal eval puts Grok 4.7 at #3 across 22 knowledge-work tasks for under $5 — realsohamparekh · 2026-09-22