Grok 4.7 tops Harvey LAB-AA v1.1, beating Claude Opus 5.5 and GPT-6 Astra
XFreeze · x · 2026-10-09
XFreeze reports that Grok 4.7 ranked #1 on the Harvey LAB-AA v1.1 leaderboard, outperforming Claude Opus 5.5, GPT-6 Astra, Muse Spark 1.3 and others. The benchmark is unusually harsh: a single material hallucination zeroes the entire task, so it measures end-to-end task completion without substantive hallucinations rather than just answer accuracy — evidence, the author argues, that Grok 4.7 is excelling at high-stakes reasoning.
More from Models
- Only 7% of training compute goes to pretraining? repligate probes what the figure means — repligate · 2026-10-10
- New GPT Memory Consolidation Model 'Memory4 Dream' Appears, Appears Built on GPT-6 Luna — lyraxana · 2026-10-10
- Grok Bot now acts as an autonomous X research analyst with daily briefings — FinanceYF5 · 2026-10-10
- Doctors Are Building Board-Style Exams for Medical AI, Starting with Radiology's Last Exam — DrDatta_AIIMS · 2026-10-10
- AI roundup: OpenAI tops 700 math papers, Mistral ships Le Chonk open model — FinanceYF5 · 2026-10-10
- Notes from an NYC AI dinner: agents find most inference perf wins, code review deemed unproductive — paulnovosad · 2026-10-10