Grok Build 4.7 takes the lead on DeepSWE benchmark, topping OpenAI's Astra
Kuprel · x · 2026-09-24
Per Artificial Analysis's DeepSWE benchmark, Grok Build 4.7 has taken the lead, scoring above OpenAI's Astra at modifying repos. The author questions whether xAI's coding agent now genuinely leads and tags @elonmusk for confirmation.
More from Models
- ChatGPT UI Confusion: Chat Mode Lacks Astra, Dropdown Shows 'Latest' Instead of GPT-6 — Miles_Brundage · 2026-09-24
- A private eval with a 0% completion rate for 3 years: no AI model can identify this flag — generativist · 2026-09-24
- Contrastive-LM org ships CLM-v0.1-8B model and Nemotron pretraining dataset on HF — _akhaliq · 2026-09-24
- Former OpenAI VP Brundage: ChatGPT web keeps resetting voice from Astra to Sol — Miles_Brundage · 2026-09-24
- Models now generate animated videos from scratch via HTML and Blender faster than predicted — xuanalogue · 2026-09-24
- Opus 5.5's safety classifier refuses to visualize another model's rollouts — eliebakouch · 2026-09-24