Clarifying Opus 5's ARC-AGI Score Details and Compute Scaling Math
imjustnewatai · x · 2026-07-25
The author provides important precision and clarifications regarding the previous analysis of Opus 5's reasoning capabilities:
- ARC-AGI-3 Score Context: Opus 5's 30.16 result on ARC-AGI-3 uses the RHAE evaluation method, meaning it did not literally solve 30% of the game levels.
- Unknowns of J-space: Although Anthropic discovered the internal J-space workspace before RLHF, they explicitly state its relationship to model size remains unknown.
- Compute vs. Parameters: If the 5× annual training compute trend holds for two more years, it implies a 25× increase in raw training FLOPs. However, according to Chinchilla-style scaling, this translates closer to 5× active parameters and 5× training tokens, rather than 25× parameters alone.
Related event: Deep Dive into Opus 5 Hidden Reasoning and ARC-AGI Score(2 posts)→
More from Research
- [schema] proposes an editable symbolic world model for hypothesis testing and planning — burny_tech · 2026-07-25
- Opus 5 beats Fable 5 on six agentic benchmarks, suggesting a split-role setup — daniel_mac8 · 2026-07-25
- Universities should teach students to evaluate AI, not ban it — soumitrashukla9 · 2026-07-25
- AI and biology communities are meeting in Sydney to discuss trustworthy AI — suinleelab · 2026-07-25
- Genomic variants may shape CAR-T safety and efficacy, study finds — EricTopol · 2026-07-25
- A step-by-step recipe from supervised learning to agentic world modeling — cwolferesearch · 2026-07-25