ARC-AGI's Kamradt: clever harnesses measure human intelligence, not models

GregKamradt · x · 2026-09-04

Greg Kamradt, who runs ARC-AGI, clarified benchmark methodology in an exchange with Gary Marcus: they knew early on you could bake human intelligence into a harness — including some wild system prompts — to do well on v3.

But at that point, he argues, you're measuring human intelligence rather than the model. That's why the team had to be more principled about which questions and tests to actually ask, so the benchmark evaluates the model rather than prompt engineering.

Original post →

More from Models

Models channel →