Grok 4.7 scores 94% on Next.js evals, 2x-7x cheaper than Opus 5.5 and GPT-6 Sol
elonmusk · x · 2026-09-23
rauchg's team ran fresh Next.js coding evals: Claude Opus 5.5, GPT-6 Sol and Fable 5.1 all scored 97%, while Grok 4.7 came in at 94% — but at 2x-7x lower cost. Elon Musk reposted, calling Grok 4.7's performance strong for a smallish model. The takeaway: frontier coding capability has converged, and price is becoming the differentiator.
More from Models
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23
- Claude Opus 5.5 costs $5.98 per task as price cuts offset ~80% token usage spike — ArtificialAnlys · 2026-09-23
- Claude Opus 5.5 benchmarked: intelligence 58, per-task cost spans 11x across five tiers — ArtificialAnlys · 2026-09-23
- GPT-6 prompt caching gets more reliable: higher hit rates, diagnostics, up to 90% savings — rohanpaul_ai · 2026-09-23
- steipete quips Anthropic 'only bans the popular' third-party coding harnesses — steipete · 2026-09-23
- Gemini training details dissected: groupwise reward redistribution to fight reward hacking — nrehiew_ · 2026-09-23