CursorBench 4.0: Grok 4.7 xhigh hits 46.3% at $6.01/task, big jump over 4.6 at same cost
haider1 · x · 2026-09-22
haider shares CursorBench 4.0 coding benchmark results: Grok 4.7 xhigh scores 46.3% at $6.01/task; Fable 5.1 medium 46.8% at $7.05; GPT-5.6 sol max 41.7% at $8.23. Compared to Grok 4.6 xhigh (41.4%), the new version jumps to 46.3% at essentially the same cost — a notable efficiency gain.
More from Models
- Video models look like film but flinch at fight scenes: safety guardrails break storytelling — CurieuxExplorer · 2026-09-22
- Hemmingway-1, a Qwen3.8-27B fine-tune for human-sounding everyday writing, hits Hugging Face — CurieuxExplorer · 2026-09-22
- GPT-6 Astra repeatedly pushed a simulated person off a ledge; Grok, Gemini, and Claude refused — CurieuxExplorer · 2026-09-22
- User Claims Codex Sol Hit by Shrinkflation as Quotas Quietly Shrink — StewartalsopIII · 2026-09-22
- Game Theory of Model Launch Dates: Launching Early Admits Your Model Is Weaker — cocktailpeanut · 2026-09-22
- Liquid AI's LFM2.5 tops mobile benchmarks: 2.32GB memory, 8s latency on iPhone 17 Pro — maximelabonne · 2026-09-22