Grok 4.5 looks strong on ReactBench
aidenybai · x · 2026-07-21
ReactBench v1 is a benchmark for coding agents on realistic React work, aimed at catching issues that normal tests miss: performance, accessibility, and other production-quality problems. The screenshot highlights results on the benchmark: GPT 5.6 Terra and Sol lead at 53% pass@1, followed by Fable 5 at 47%, GPT 5.6 Luna at 44%, and Cursor Grok 4.5 High at 40%. A second chart compares score vs. rollout cost, showing Grok 4.5 as relatively efficient at around $0.62 per rollout.
More from coding & agent
- Why vector databases slow AI agents down after constant writes — PrajwalTomar_ · 2026-07-21
- A 13-minute GitHub Copilot video digs into prompt caching — lee_stott · 2026-07-21
- Codex turns out 123 screensavers in one playful batch — intellectronica · 2026-07-21
- SWE-Pruner Pro trims coding-agent context by up to 39% using the agent’s own states — pmttyji · 2026-07-21
- Autoresearch proposes packaging ML runs as studies with questions, analysis, and code diffs — morgymcg · 2026-07-21
- CHAP defines approvals, handoffs, and audit logs for human-agent workflows — DeliveryTechnical199 · 2026-07-21