ARC-AGI's Kamradt: clever harnesses measure human intelligence, not models
GregKamradt · x · 2026-09-04
Greg Kamradt, who runs ARC-AGI, clarified benchmark methodology in an exchange with Gary Marcus: they knew early on you could bake human intelligence into a harness — including some wild system prompts — to do well on v3.
But at that point, he argues, you're measuring human intelligence rather than the model. That's why the team had to be more principled about which questions and tests to actually ask, so the benchmark evaluates the model rather than prompt engineering.
More from Models
- Databricks evals: GPT-6 Astra claims SOTA on OfficeQA Pro benchmarks, cheaper per task — downingARK · 2026-09-04
- Mathematician tests GPT-6 Astra: live Lean proof verification while writing arguments — teortaxesTex · 2026-09-04
- Tavus Launches Sparrow-2, Claiming #1 in End-of-Turn Detection and Interruption Handling — ycombinator · 2026-09-04
- ARC-AGI-3: 10x reasoning tokens cuts total cost from $48k to $26k vs medium — i_dg23 · 2026-09-04
- Researchers flag data contamination concerns in benchmark behind Astra's time-horizon score — dfrsrchtwts · 2026-09-04
- antirez: judge new models by whether they fix real blocking bugs, not three.js demos — antirez · 2026-09-04