Opus 5, hidden-rule inference, and J-space point to a new agent stack
imjustnewatai · x · 2026-07-25
The post argues that a line in the Opus 5 system card hints at how future agents may infer hidden structure before acting. It connects several pieces of evidence: an ARC-style run where Opus 5 derived a reflection equation, solved all 8 levels in 294 actions, and then generated 60 mirrored targets before acting; a preprint finding that hidden-rule inference rises sharply with model size across 11 LLMs from 0.5B to 72B parameters; and Anthropic’s separate discovery of a small causal workspace called J-space, which still preserves fluency when present but collapses multistep reasoning when suppressed.
The broader thesis is that the frontier stack is becoming: pretraining builds the workspace, scale strengthens hidden-rule inference, post-training teaches executable abstractions, and compute expands both training and test-time search. The author also extrapolates a compute runway of roughly 5× per year, mentions Anthropic’s planned 1GW deployment starting in 2027, and cites a possible ≥10GW OpenAI + NVIDIA buildout. The conclusion: by 2027–2028, agents may enter a world, infer its hidden variables, write the rules, test them, and act while learning the game.
Related event: Deep Dive into Opus 5 Hidden Reasoning and ARC-AGI Score(2 posts)→
More from Models
- Opus 5 reportedly scores 42/42 on IMO 2026 without tools — Afinetheorem · 2026-07-25
- Quadrillion says Anthropic’s Opus 5 is faster than Opus 4.8 on hard ML workloads — igarciacamargo · 2026-07-25
- Google is lagging behind open-weight models on most benchmarks — burny_tech · 2026-07-25
- Claude Opus 5 tops an Artificial Analysis coding-agent benchmark at 67 — Hesamation · 2026-07-25
- Side-by-side eval shows diffusion loses overall, but wins speed in agent loops — Additional-Engine402 · 2026-07-25
- Charts Show Opus 5 Peaks in Coding Performance with Medium Thinking — dejavucoder · 2026-07-25