Grok 4.5 tops HighWalk benchmark for updating specs from code changes
elonmusk · x · 2026-07-29
Grok 4.5 (high) is shown topping the HighWalk benchmark, which measures how well AI agents update technical specifications from code changes.
The chart shared in the post says Grok beat Claude and GPT on a combination of quality and operational efficiency.
More from Models
- Scobleizer says Grok 4.5 is the best coding model right now — Scobleizer · 2026-07-29
- Apple reportedly sues OpenAI over alleged trade-secret misuse tied to future hardware — emmanuelvivier · 2026-07-29
- Kimi K3 and Qwen3.8 show Chinese AI is now a sustained competitive force — emmanuelvivier · 2026-07-29
- GPT-5.6 reportedly solves Feige’s 1/e conjecture in probability — MickeySteamboat · 2026-07-29
- Kimi K3 Q2 reportedly runs locally on two 512GB M3 Ultra Mac Studios — pcuenq · 2026-07-29
- Gemini Hilariously Overthinks Math Riddle as a 9/11 Dark Joke — TourPsychological800 · 2026-07-29