Grok 4.7 jumps to 46.3% on CursorBench and 64% on EEBench, keeping the same $2/$6 per million token pricing
FinanceYF5 · x · 2026-09-22
Comparing open-world city game demos built by Grok 4.7 vs 4.6, the new model works longer, verifies more, and builds better.
Benchmarks:
- CursorBench: 40.4% → 46.3%
- EEBench: 53% → 64%
Pricing unchanged: $2/M input tokens, $6/M output tokens.
More from coding & agent
- Google open-sources ARTEMIS, an AI agent that drives real Android phones at 3s per step — aigclink · 2026-09-22
- Google open-sources ARTEMIS, an on-device Android GUI agent at 3s per step — aigclink · 2026-09-22
- vackrooms: an open-source endless backrooms game built entirely with Claude — pablostanley · 2026-09-22
- Parallel beats serial at equal budget: best-of-N@T far outperforms best-of-1@N*T — ShangyinT · 2026-09-22
- Give Your LLM a Wiki: demo cuts input tokens by ~50% — kedar5 · 2026-09-22
- Claude Code Projects called the best interface yet for parallel serious work — daniel_mac8 · 2026-09-22