Grok 4.5 Programming Benchmark Comparison

jasonkneen · x · 2026-07-11

A programming comparison was reposted: using the same prompt, the results from running Grok 4.5 via the Grok CLI harness were "quite good" and more cost-effective.

The quoted content explains that the test task was to build a three.js procedural terrain generator, comparing GPT-5.6 Sol Ultra, Claude Fable 5, and results using the Codex / Claude Code harness. The author also emphasized that no agent skills were invoked, and the prompt can be viewed in the original post.

Related event: Frontier Models Compete on Visual Coding Tasks(4 posts)→

Original post →

More from coding & agent

coding & agent channel →