High Scores Reported on ARC-AGI-3
xiuyu_l · x · 2026-07-17
This post shares a new result regarding ARC-AGI-3: a harness named schema has achieved exceptionally high scores on the public set.
Key metrics include:
- 99% RHAE using Opus 4.8 + Fable 5
- 95.35% RHAE using GPT-5.6 Sol
- The author claims that schema enables LLMs to "think like physicists"
Reposts and comments emphasize that these results indicate programmable world models could significantly boost model and agent performance. Furthermore, with the advent of stronger agentic models, symbolic discovery may become a crucial pathway for problem-solving.
Related event: Schema Harness Sparks ARC-AGI-3 Debate(14 posts)→
More from Models
- Unreleased 'GPT 6 Sol' model spotted in OpenAI's API — SteveEricJordan · 2026-09-11
- DeepSeek update keeps cache hits mid-conversation, cuts costs 36.6% — teortaxesTex · 2026-09-11
- AI Sextet offers 6 models free and unlimited for 14 days, including DeepSeek and Qwen — airesearch12 · 2026-09-11
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11