OpenAI's Astra Scored 62.7% and 99.9% on the Same Benchmark, 37 Points Apart

mixtapedmonk · reddit · 2026-09-14

Digging into ARC Prize's actual results table, the author found OpenAI's GPT-6 Astra posted 62.7% and 99.9% on the same ARC-AGI-3 benchmark depending on which harness was used—37 points apart—and the org that built the test won't call it AGI. Fortune separately found five numbers quietly changed on OpenAI's own launch page after it went live. Comparing the Llama 4 benchmark precedent, the writeup concludes it could be genuine harness noise or something else, but headline-screenshot benchmark reporting is clearly broken.

Original post →

More from Models

Models channel →