ARC Prize: Harness, Not Model, Drives GPT-6 Astra Scores
ARC Prize showed GPT-6 Astra scores on ARC-AGI-3 swing from 63% to 97-99.9% depending on the harness used, arguing that scaffolding code can matter more than the model itself.
2026-09-07 ~ 2026-09-08 · 2 related posts
- Same model, 63% vs 97% on ARC-AGI-3: 7 harness rules beat max reasoning — victor_explore · 2026-09-07
- ARC Prize: OpenAI's 99.9% AGI score came from its harness, not the model — GaryMarcus · 2026-09-08