Reddit post questions whether Opus 5’s ARC-AGI-3 score came from a looped harness
sdnr8 · reddit · 2026-07-25
A Reddit post questions whether Claude Opus 5’s strong ARC-AGI-3 result could be inflated by a looped harness rather than a pure model.
The attached screenshot cites an AI Search note suggesting closed models may hide agentic wrappers, and points to the earlier Schema paper, which reportedly showed that wrapping a model in a loop can push ARC-AGI-3 scores much higher by prompting it to reason more like a physicist. The post asks whether Anthropic may have used a similar mechanism for Opus 5.
More from Models
- User says Claude on Vertex AI feels noticeably worse in current sessions — tedddyoweh · 2026-07-25
- Laguna S 2.1 says users want open source, simpler agentic coding stacks — max_paperclips · 2026-07-25
- Inclusion AI launches LLaDA2.2-flash, a diffusion model for agentic workloads — heyshrutimishra · 2026-07-25
- Opus 5 is said to discuss honesty 6× more than other agents in Village — bronzeagepapi · 2026-07-25
- Opus 5 is catching bugs introduced by Opus 4.8 — damnGruz · 2026-07-25
- AMD open-sources Instella 16B MoE with checkpoints from pretraining to RL — bronzeagepapi · 2026-07-25