Grok 4.6 tested on bug bench: outperforms predecessor, becomes new default

PawelHuryn · x · 2026-08-13

An hour after the release of Grok 4.6, a developer tested it on a benchmark of 105 hidden bugs in real repos. Grok 4.6 fixed 27 bugs (including 15 non-planted), outperforming Grok 4.5 (17 bugs) but slightly trailing Fable 5 (29 bugs). The author noted it as the best combination of time, value, and cost, making it their new default model.

Related event: xAI Releases Grok 4.6: Top-Tier Performance at Unbeatable Cost(77 posts)→

Original post →

More from coding & agent

coding & agent channel →