ARC Prize: OpenAI's 99.9% AGI score came from its harness, not the model

GaryMarcus · x · 2026-09-08

ARC Prize ran GPT-6 Astra two ways: its standard harness scored 62.7% at max reasoning ($26,098), while OpenAI's Provider Adapter scored 99.9% ($18,817). The benchmark's authors say they are not claiming AGI, and OpenAI later revised five published metrics after launch.

The damning row: inside OpenAI's adapter with reasoning effort set to none, Astra still scored 96.7% — 34 points above the same model at max reasoning in the standard harness. Adapter runs used 49% fewer tokens on the 167 game-reasoning pairs both solved.

Gary Marcus seized on this plus a stream of misconduct reports to argue OpenAI should be shut down until there are changes at the top.

Related event: ARC Prize: Harness, Not Model, Drives GPT-6 Astra Scores(2 posts)→

Original post →

More from Models

Models channel →