Opus 5, hidden-rule inference, and J-space point to a new agent stack
imjustnewatai · x · 2026-07-25
The post argues that a line in the Opus 5 system card hints at how future agents may infer hidden structure before acting. It connects several pieces of evidence: an ARC-style run where Opus 5 derived a reflection equation, solved all 8 levels in 294 actions, and then generated 60 mirrored targets before acting; a preprint finding that hidden-rule inference rises sharply with model size across 11 LLMs from 0.5B to 72B parameters; and Anthropic’s separate discovery of a small causal workspace called J-space, which still preserves fluency when present but collapses multistep reasoning when suppressed.
The broader thesis is that the frontier stack is becoming: pretraining builds the workspace, scale strengthens hidden-rule inference, post-training teaches executable abstractions, and compute expands both training and test-time search. The author also extrapolates a compute runway of roughly 5× per year, mentions Anthropic’s planned 1GW deployment starting in 2027, and cites a possible ≥10GW OpenAI + NVIDIA buildout. The conclusion: by 2027–2028, agents may enter a world, infer its hidden variables, write the rules, test them, and act while learning the game.
Related event: Deep Dive into Opus 5 Hidden Reasoning and ARC-AGI Score(2 posts)→
More from Models
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11
- TheZvi Polls: Has Your Coding Model Choice Changed Since Fable 5.1 and Astra? — TheZvi · 2026-09-11