Reddit post questions whether Opus 5’s ARC-AGI-3 score came from a looped harness

sdnr8 · reddit · 2026-07-25

A Reddit post questions whether Claude Opus 5’s strong ARC-AGI-3 result could be inflated by a looped harness rather than a pure model.

The attached screenshot cites an AI Search note suggesting closed models may hide agentic wrappers, and points to the earlier Schema paper, which reportedly showed that wrapping a model in a loop can push ARC-AGI-3 scores much higher by prompting it to reason more like a physicist. The post asks whether Anthropic may have used a similar mechanism for Opus 5.

Related event: Opus 5's High ARC-AGI-3 Score Sparks Cheating Allegations and Benchmark Debate(8 posts)→

Original post →

More from Models

Models channel →