Chart shows ARC-AGI-1 is now officially harder for models than ARC-AGI-3
burny_tech · x · 2026-09-04
FakePsyho shared a chart declaring it now official: ARC-AGI-1 is harder than ARC-AGI-3. The counterintuitive takeaway is that the newer ARC-AGI-3's task design makes it more tractable for reasoning models than the original benchmark, sparking discussion about benchmark validity.
More from Research
- Martian says routing across 44 LLMs cuts errors 46% vs best single model on 16 benchmarks — Arindam_1729 · 2026-09-04
- Bengio co-authors arXiv framework for monitoring rogue AI progression, born from FFRDC cross-lab workshop — Miles_Brundage · 2026-09-04
- FlashRender: Few-step camera-controlled generative rendering via MeanFlow distillation — everex · 2026-09-04
- LatentStream: progressive latent memory evolution for streaming video understanding — Hongyu Qu · 2026-09-04
- Temporal Context Routing aligns script timing in joint audio-video generation — Yichen Liu · 2026-09-04
- Puffin-World scales unified multimodal model with native 3D world states — Kang Liao · 2026-09-04