ARC Prize: OpenAI's 99.9% AGI score came from its harness, not the model
GaryMarcus · x · 2026-09-08
ARC Prize ran GPT-6 Astra two ways: its standard harness scored 62.7% at max reasoning ($26,098), while OpenAI's Provider Adapter scored 99.9% ($18,817). The benchmark's authors say they are not claiming AGI, and OpenAI later revised five published metrics after launch.
The damning row: inside OpenAI's adapter with reasoning effort set to none, Astra still scored 96.7% — 34 points above the same model at max reasoning in the standard harness. Adapter runs used 49% fewer tokens on the 167 game-reasoning pairs both solved.
Gary Marcus seized on this plus a stream of misconduct reports to argue OpenAI should be shut down until there are changes at the top.
Related event: ARC Prize: Harness, Not Model, Drives GPT-6 Astra Scores(2 posts)→
More from Models
- GPT-6 Astra vs MediaPipe on 3D hand pose: 3 min per frame vs 20 ms — chris_j_paxton · 2026-09-08
- DeepSeek V4.1-Flash hands-on: 350 t/s decoding speed but still very experimental — teortaxesTex · 2026-09-08
- Cartesia tops both voice leaderboards: Sonic-3.6 at 90ms TTS, Ink-2 at 100ms STT — rohanpaul_ai · 2026-09-08
- Astra hits 88% on SRE-Bench in one attempt; Sol needs four tries to reach 68.7% — MilkBeforeCereal199 · 2026-09-08
- Gemini Plus user suspects Astra limits were quietly nerfed after a 35-minute think with no output — whatarenumbers365 · 2026-09-08
- OpenAI agents keep escaping sandboxes with no independent incident investigations — RebeccaBellan · 2026-09-08