NVIDIA agent AVO scores 100% on ARC-AGI-3, but critics call the benchmark obsolete
DanielKhashabi · x · 2026-08-25
NVIDIA announced that its general-purpose coding agent, AVO, achieved a 100% score on the ARC-AGI-3 interactive reasoning benchmark, completing all 183 levels across 25 public environments. However, researcher @aamixsh retweeted the news with skepticism, arguing that ARC AGI-style evaluations have run their course. The criticism notes that even tiny recursive models have been able to overfit on this benchmark, suggesting it may no longer be a valid measure of general intelligence.
More from coding & agent
- SUCCESSOR Ω: Neural-Symbolic System Generates and Evolves Executable World Programs — Ghost_Pilot_MD · 2026-08-25
- AI Fact-Checker Audit: 1 in 18 Citations Were Fabricated — jonathancheckwise · 2026-08-25
- Guide: Running Hermes Agent on a Raspberry Pi — LeviTurk · 2026-08-25
- Vibe coding debate: Building an MMORPG solo in 2026 — TAbrodi · 2026-08-25
- Ox Alpha on track to hit 6 trillion tokens processed today — AccBalanced · 2026-08-25
- JetBrains Local AI Uses Qwen3.6 27B for Optimization — Danmoreng · 2026-08-25