Grok 4.5 looks strong on ReactBench

aidenybai · x · 2026-07-21

ReactBench v1 is a benchmark for coding agents on realistic React work, aimed at catching issues that normal tests miss: performance, accessibility, and other production-quality problems. The screenshot highlights results on the benchmark: GPT 5.6 Terra and Sol lead at 53% pass@1, followed by Fable 5 at 47%, GPT 5.6 Luna at 44%, and Cursor Grok 4.5 High at 40%. A second chart compares score vs. rollout cost, showing Grok 4.5 as relatively efficient at around $0.62 per rollout.

Original post →

More from coding & agent

coding & agent channel →