ARC-AGI-4 should stay private, after Opus 5 scored 3× the next-best model on ARC-AGI-3
burny_tech · x · 2026-07-25
- The post argues that ARC-AGI-4 should probably be kept more private, with far fewer public demos.
- It suggests that many people now assume ARC-AGI progress is being driven by targeted RL environments rather than broad, public benchmark play.
- The quoted context notes that on ARC-AGI-3, a benchmark where models must solve novel problems, Opus 5 scored 3× higher than the next-best model.
- The discussion is less about a product announcement and more about how benchmark design and publicity may affect what progress looks like.
More from Models
- FrontierCode’s design explains why Opus 5 scores drop as reasoning increases — silasalberti · 2026-07-25
- Researcher claims a universal jailbreak works across GPT-5.6, Opus 5 and Fable — Polymarket · 2026-07-25
- Users say Claude and Gemini are removing reasoning summaries they liked — breath_mirror · 2026-07-25
- GPT-6 could land within two weeks, with monthly model updates becoming normal — haider1 · 2026-07-25
- User Praises Claude Opus 5 as the Best Model for Daily Work — rudrank · 2026-07-25
- Anthropic’s system cards may finally be getting shorter, says Miles Brundage — Miles_Brundage · 2026-07-25