Grok 4.5 Ties Competitor on Coding Benchmark
elonmusk · x · 2026-07-11
Grok 4.5 paired with Grok Build has tied with Codex GPT-5.6 on the SWE-Atlas-QnA benchmark, scoring 84. The repost describes this as another leap forward for xAI in coding and agentic capabilities, noting that this feature is now accessible via the new Meta Model API and Meta AI.
Related event: Grok 4.5 Tops Coding Benchmark with High Token Efficiency(3 posts)→
More from coding & agent
- Why 88-95% of enterprise AI agent pilots never ship — and what working teams do differently — ankitsharma112 · 2026-09-07
- Open-source Graft fights coding agent amnesia with markdown, claims SWE-bench win over Claude Code — thisdudelikesAI · 2026-09-07
- New Agent Workflow: Have AI Implement a Feature Once to Learn, Then Rebuild From Scratch — remilouf · 2026-09-07
- Founder's real-time AI avatar handled inbound sales during paternity leave, closing prospects — toolstelegraph · 2026-09-07
- Dev mocked as vibe coder claps back: I can write FizzBuzz in under 15 minutes — tlakomy · 2026-09-07
- Running 4 parallel agents feels like babysitting 4 toddlers, dev says — smlpth · 2026-09-07