NVIDIA's AVO Agent Scores 100% on ARC-AGI-3 Public Set
On August 21, NVIDIA released AVO, a general-purpose coding agent that scored 100% on the ARC-AGI-3 interactive reasoning benchmark, completing all 183 levels across the 25 public environments. This is a rare perfect run on the benchmark, and according to @eigenhector, who reshared the news, AVO was not specifically designed for the benchmark—it accomplishes tasks with minimal inputs and tools, simply by examining past grids represented as text.
Confirmed
- AVO scored 100% on ARC-AGI-3, completing 183 levels across 25 environments, a result confirmed by multiple posts from @theologi, @NVIDIAAI, and @MagicZhang
- AVO's architecture is named Agentic Variation Operators, built by integrating persistent memory, supervised feedback, and tool use; according to @NVIDIAAI's post, it incorporates the capabilities of models like Claude into the system design
- Without any instructions, explicit rules, or predefined goals, AVO successfully inferred task objectives and solved them, demonstrating strong generalization in reasoning
- @MagicZhang shared a screenshot of the official scorecard showing the 100% score
Why it matters
- In its release, @NVIDIAAI emphasized that AVO's result suggests system design (agent architecture, memory, and tool integration) may matter more than simply scaling up model capabilities, pointing to a new direction for general-purpose agent research
- Achieving a perfect score without benchmark-specific tuning signals that the system's autonomous goal inference and generalization in open environments deserve attention
2026-08-21 ~ 2026-08-22 · 18 related posts
Primary sources
- NVIDIA AVO agent scores 100% on ARC-AGI-3 benchmark — NVIDIAAI ·
- NVIDIA AVO achieves 100% on ARC-AGI benchmark — theologi ·
- NVIDIA's AVO harness lifts Opus 5 from 30% to 100% on ARC-AGI-3 — daniel_mac8 ·
- [source] NVIDIA AVO agent scores 100% on ARC-AGI-3 benchmark — NVIDIAAI · 2026-08-21
- NVIDIA AVO achieves 100% on ARC-AGI-3, proving system design over model capability — NVIDIAAI · 2026-08-21
- [source] NVIDIA AVO achieves 100% on ARC-AGI benchmark — theologi · 2026-08-21
- NVIDIA's Coding Agent Scores 100% on ARC-AGI-3 Benchmark — MagicZhang · 2026-08-21
- NVIDIA's AVO Agent Scores 100% on ARC-AGI-3 Benchmark — eigenhector · 2026-08-21
- [source] NVIDIA's AVO harness lifts Opus 5 from 30% to 100% on ARC-AGI-3 — daniel_mac8 · 2026-08-22
- Nvidia AVO Lacks Open Code, Similar to Open Source Prime Agent — daniel_mac8 · 2026-08-22
- NVIDIA's AVO Achieves 100% on ARC-AGI-3, Outperforms FlashAttention-4 — daniel_mac8 · 2026-08-22
- No open-source for NVIDIA's AVO? Prime Intellect's Prime Agent is a close cousin — daniel_mac8 · 2026-08-22
- Nvidia paired Claude Opus 5 with memory and a supervisor to score 100% on ARC-AGI-3 — HaktanSuren · 2026-08-22
- NVIDIA hits 100% on ARC-AGI-3 public set via agent harness — mhmazur · 2026-08-22
- NVIDIA's AVO framework boosts Claude Opus to 100% on ARC-AGI benchmark — imjustnewatai · 2026-08-22
- NVIDIA releases AVO agent framework; Chollet clarifies 100% demo score is not benchmark pass — jamestagg · 2026-08-22
- NVIDIA's AVO agent scores 100% on ARC-AGI-3, beats FlashAttention-4 in kernel optimization — MickeySteamboat · 2026-08-22
4 near-duplicate retellings: wavefnx · JFPuget · daniel_mac8 · ZeroStateReflex