NVIDIA agent AVO scores 100% on ARC-AGI-3, but critics call the benchmark obsolete

DanielKhashabi · x · 2026-08-25

NVIDIA announced that its general-purpose coding agent, AVO, achieved a 100% score on the ARC-AGI-3 interactive reasoning benchmark, completing all 183 levels across 25 public environments. However, researcher @aamixsh retweeted the news with skepticism, arguing that ARC AGI-style evaluations have run their course. The criticism notes that even tiny recursive models have been able to overfit on this benchmark, suggesting it may no longer be a valid measure of general intelligence.

Original post →

More from coding & agent

coding & agent channel →