Dev Critiques ARC-AGI-3: Already Saturated, Quadratic Scoring Inflates Egos

mgostIH · x · 2026-07-30

A developer points out that the ARC-AGI-3 benchmark is already saturated by current testing harnesses, meaning even minor reinforcement learning (RL) can massively inflate model scores.

Furthermore, the author heavily criticizes the benchmark's scoring system. They argue that calculating the score using a quadratic percentage of solved tasks serves no purpose other than to artificially boost certain egos, especially when early scores were all near 0%.

Related event: ARC-AGI 3 Benchmark Criticized for Flawed Scoring and Memory Rules(3 posts)→

Original post →

More from Research

Research channel →