Grok 4.7 beats 4.6 by just 0.1 across runs: 'random noise,' independent eval finds

PawelHuryn · x · 2026-09-22

Pawel Huryn's multi-run evals found Grok 4.7 (max) beat 4.6 (max) by just 0.1 points on n=4, calling it random noise and stopping that effort level. At xhigh: 4.6 scored 27/30/29, 4.7 scored 25/30/31 across three runs — none beating the median of Muse Spark 1.3. Tests used hard problems frontier models missed in early 2026, not planted bugs.

Original post →

More from Models

Models channel →