Claude Opus 5 + Claude Code + 1 Skill Scores 100% on ARC AGI 3
Tolopono · reddit · 2026-08-20
A blog post demonstrates that using Claude Opus 5 alongside Claude Code and a specific skill results in a 100% score on the public set of the ARC-AGI 3 benchmark. The author suggests that this might indicate the benchmark is not as hard as previously thought, given the right configuration.
More from Models
- The 'alignment tax': corporate AI guardrails add 25-35% to compute bills — vasilisvj · 2026-08-20
- DeepSeek: Fast-growing AI project known for reasoning and cost-efficiency — goyalshaliniuk · 2026-08-20
- Qwen 3: Alibaba's powerful open-source LLM excels in reasoning and coding — goyalshaliniuk · 2026-08-20
- MiniMax H3 measured: BGM vanishes at 480p; clip length buys it back — nakamep · 2026-08-20
- Claude fails badly at cat castle design vs Gemini and GPT — Alarming-Koala-3524 · 2026-08-20
- DeepSeek vs Opus 5 Routing Strategy and Performance Debate — teortaxesTex · 2026-08-20