Inkling-Small Sets New Open-Weight Record on ARC-AGI Benchmark
tessybarton · x · 2026-07-31
Thinking Machines' Inkling-Small has achieved record-breaking results for open-weight models on the ARC-AGI benchmark.
- ARC-AGI-2: Scored 40.1% at $0.23/task.
- ARC-AGI-1: Scored 84% at $0.11/task.
It is the highest-scoring open-weight model evaluated by ARC Prize on both benchmarks, setting a new cost-performance frontier.
More from Models
- Anthropic Reports Claude Outages Due to Multiple Network Failures — trq212 · 2026-07-31
- Users Report Claude Opus Frequently Lies Badly in Interactions — rickasaurus · 2026-07-31
- OpenAI Accused of Cherry-Picking Data in Benchmark Graphs — ns123abc · 2026-07-31
- Agent Arena Leaderboard Updates: Claude Fable 5 Takes #1 — arena · 2026-07-31
- Sol-5.6 Ultra Mode Reported to Overthink and Get Stuck in Loops — AIandDesign · 2026-07-31
- Claude API Retains Thinking Blocks by Default; Developers Urge OpenAI to Follow Suit — steipete · 2026-07-31