Grok 4.5 tops HighWalk benchmark for updating specs from code changes

elonmusk · x · 2026-07-29

Grok 4.5 (high) is shown topping the HighWalk benchmark, which measures how well AI agents update technical specifications from code changes.

The chart shared in the post says Grok beat Claude and GPT on a combination of quality and operational efficiency.

Original post →

More from Models

Models channel →