Grok 4.6 tested on bug bench: outperforms predecessor, becomes new default
PawelHuryn · x · 2026-08-13
An hour after the release of Grok 4.6, a developer tested it on a benchmark of 105 hidden bugs in real repos. Grok 4.6 fixed 27 bugs (including 15 non-planted), outperforming Grok 4.5 (17 bugs) but slightly trailing Fable 5 (29 bugs). The author noted it as the best combination of time, value, and cost, making it their new default model.
Related event: xAI Releases Grok 4.6: Top-Tier Performance at Unbeatable Cost(77 posts)→
More from coding & agent
- Open-Source Zero Trust Platform Octelium Supports Building MCP Gateways — tom_doerr · 2026-08-13
- Showcasing Hermes Agent Desktop Plugins: Custom Builds and Interactions — Teknium · 2026-08-13
- Xcode 27 beta 5 introduces 3 new agent skills for Siri integration — rudrank · 2026-08-13
- Kimi Code 0.36.0: Main Agent Can Now Dynamically Dispatch Sub-Model Pools — KimiDevs · 2026-08-13
- Kimi Code Update Adds Full-Screen TUI Mode and LaTeX Formula Rendering — KimiDevs · 2026-08-13
- iOS 27 New Siri Dev Guide: How to Integrate with App Entities — rudrank · 2026-08-13