High Scores Reported on ARC-AGI-3
xiuyu_l · x · 2026-07-17
This post shares a new result regarding ARC-AGI-3: a harness named schema has achieved exceptionally high scores on the public set.
Key metrics include:
- 99% RHAE using Opus 4.8 + Fable 5
- 95.35% RHAE using GPT-5.6 Sol
- The author claims that schema enables LLMs to "think like physicists"
Reposts and comments emphasize that these results indicate programmable world models could significantly boost model and agent performance. Furthermore, with the advent of stronger agentic models, symbolic discovery may become a crucial pathway for problem-solving.
Related event: Schema Harness Sparks ARC-AGI-3 Debate(14 posts)→
More from Models
- NVIDIA says Nemotron 3 Ultra scored 30/42 on the 2026 IMO problems — NVIDIAAI · 2026-07-22
- OpenAI is reportedly briefing U.S. lawmakers on its next model family — kimmonismus · 2026-07-22
- Muse Spark 1.1 lands at 1495 on Text Arena with standout agentic-coding price performance — ycombinator · 2026-07-22
- Advanced AI Models Are Becoming Impossible to Plug and Play — emollick · 2026-07-22
- Google Gemini's AI Problem: No Leading Model for Core Workloads — bindureddy · 2026-07-22
- Model Offers 1M Token Context Window at Just $0.33/1M Tokens — MickeySteamboat · 2026-07-22