Don't fall for GPT-6's 98.6% ARC AGI-3: Nvidia AVO already hit 100%
Informal-Trouble2183 · reddit · 2026-09-04
Reacting to claims that GPT-6 Astra scored 98.6% on ARC AGI-3, the poster points out Nvidia already demonstrated 100% on the benchmark with its AVO harness, a frontier general-purpose architecture for long-horizon autonomous agents. Crucially, OpenAI used its own harness rather than the standard one, so the scores aren't directly comparable.
More from Models
- ChatGPT's Reddit-flavored reply to a sexual assault victim sparks training-data backlash — airkatakana · 2026-09-04
- Astra Shows Near-Ideal Test-Time Scaling Gains on LifeSciBench Benchmark — soleio · 2026-09-04
- Astra reportedly hits 97% on ARC-AGI-3 without chain-of-thought — FeeAvailable3770 · 2026-09-04
- GPT-6 'Astra' Smashes ARC-AGI Records: 62.7% on Standard Harness, 99.9% on ARC-AGI-3 — repligate · 2026-09-04
- Browser QA harness: GLM 5.3 Flash beats DeepSeek V4 Flash Vision on screenshots — Certain_Pension6305 · 2026-09-04
- Reddit user flags Artificial Analysis as unreliable: same model shows conflicting scores — metigue · 2026-09-04