Clarifying Opus 5's ARC-AGI Score Details and Compute Scaling Math
imjustnewatai · x · 2026-07-25
The author provides important precision and clarifications regarding the previous analysis of Opus 5's reasoning capabilities:
- ARC-AGI-3 Score Context: Opus 5's 30.16 result on ARC-AGI-3 uses the RHAE evaluation method, meaning it did not literally solve 30% of the game levels.
- Unknowns of J-space: Although Anthropic discovered the internal J-space workspace before RLHF, they explicitly state its relationship to model size remains unknown.
- Compute vs. Parameters: If the 5× annual training compute trend holds for two more years, it implies a 25× increase in raw training FLOPs. However, according to Chinchilla-style scaling, this translates closer to 5× active parameters and 5× training tokens, rather than 25× parameters alone.
Related event: Deep Dive into Opus 5 Hidden Reasoning and ARC-AGI Score(2 posts)→
More from Research
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11
- Fruit fly brain as an LLM: connectome-driven language model demo goes live — ngxson · 2026-09-11
- Harry Collins: LLMs can't do frontier science because they can't invent new language — whoamisri · 2026-09-11