Grok 4.7 beats GPT-6 Astra on professional work, leading by 153 Elo on GDPval-AA
XFreeze · x · 2026-09-22
On the same Artificial Analysis run, Grok 4.7 leads GPT-6 Astra by +153 Elo on GDPval-AA (professional deliverables) and +88 Elo on AA-Briefcase (multi-hour office work), putting it ahead on both professional output and long-form knowledge work. The poster says SpaceXAI is pushing Grok hard into real-world knowledge work.
More from Models
- EEBench comparison sparks debate: Grok 4.7 called out vs Astra's speed and cost — teortaxesTex · 2026-09-22
- Leak: OpenAI's new agent reportedly named Aeon, launch expected Thursday to rival Grok Bot — ZeroStateReflex · 2026-09-22
- Real telemetry contradicts Grok 4.7 ragebait: 46% fewer tokens per task — ns123abc · 2026-09-22
- Studying LLM psychology today is like psychology in 1850, researcher argues — repligate · 2026-09-22
- Meta's SAM 3.1 segmentation model spotted, used for GIF creation demos — Necessary-Garlic-704 · 2026-09-22
- Laya's First Model Went from Training to Open-Source Release in Just 15 Hours — NirantK · 2026-09-22