Grok 4.7 beats 4.6 by just 0.1 across runs: 'random noise,' independent eval finds
PawelHuryn · x · 2026-09-22
Pawel Huryn's multi-run evals found Grok 4.7 (max) beat 4.6 (max) by just 0.1 points on n=4, calling it random noise and stopping that effort level. At xhigh: 4.6 scored 27/30/29, 4.7 scored 25/30/31 across three runs — none beating the median of Muse Spark 1.3. Tests used hard problems frontier models missed in early 2026, not planted bugs.
More from Models
- Xiaomi's MiMo-V2.6-Pro debuts as top open-weights model with 46 on AA Intelligence Index — huggingface · 2026-09-22
- Open multilingual System 1 decision model tops Hugging Face trending — huggingface · 2026-09-22
- Independent Tests Show Grok 4.6 (high) Beating 4.7, 23 vs. 19 — PawelHuryn · 2026-09-22
- mimo-v2.6-pro Claimed to Redraw the Price-Performance Pareto Frontier — zainhas · 2026-09-22
- Unverified: Grok 4.7 out now, costs more per task than GPT, says leaker — ChrisGPT · 2026-09-22
- Using a Haiku subagent to tame Opus 5's verbose outputs — WasdAcid · 2026-09-22