PrimeAgent faces backlash: ARC-AGI-3 high score accused of overfitting

机器之心 · wechat · 2026-08-07

Open-source agent framework PrimeAgent by PrimeIntellect claims 95.5% on ARC-AGI-3 with Opus 5, surpassing humans, but its design and validity sparked intense community debate.

Core Design

Backlash

Peter Wang, co-founder of Shortcut AI, raised two core challenges:

Author's Response

Alex Zhang, lead author of the RLM paper, countered that infinite recursion isn't RLM's core, and a depth of 1 doesn't mean weak expressiveness. The limit is merely for cost control. He admitted ARC-AGI-3 is an exploitable benchmark but emphasized the ContinualHarness design as the true differentiator.

Related event: PrimeIntellect open-sources Prime Agent, scores 95.5% on ARC-AGI-3, beating human experts(40 posts)→

Original post →

More from coding & agent

coding & agent channel →