Kepler hits server-verified 100 on all 25 ARC-AGI-3 games with Opus 5, for $777.72
rohanpaul_ai · x · 2026-10-04
The arXiv paper "Kepler: Auditable World Models for ARC-AGI-3" (Wensen Wu, NeurIPS 2026 workshop) presents an open-source harness that represents hypotheses as executable world models validated via retrospective transition checks and conditional prediction checks. Under one frozen Claude Opus 5 configuration, Kepler obtained a server-verified 100.00 RHAE on all 25 public games with no per-game model selection; in 181 of 183 completed levels its first attempt used no more actions than the median-human baseline. Total cost: 858M tokens, 97.37% cache reads, $777.72. On the final Opus 5 and GPT-5.6 boards, 48 of 50 game-model cells hit 100.
The paper also reports three evaluation failures: source-code leakage producing an invalid perfect run, agents reconstructing a removed harness in a control condition, and autonomous repair masking a broken planner. A case study found animation frames contained task-relevant information absent from text grids. Conclusion: public-set scores alone have limited discriminative value; first-attempt, cost-conditioned, verification-aware reporting is needed.
More from coding & agent
- Grok bot ran codex autonomously for 10 hours straight without asking questions — mazzaTalk · 2026-10-04
- Ramen 0.6.0: Self-Hosted Multi-Zone MCP Server for GKE/EKS with OAuth — Ok_Plum3595 · 2026-10-04
- Open-source Codex proxy setup plugs locally run Gemma models in seamlessly — TheZachMueller · 2026-10-04
- What Breaks When Your AI Agent Browses the Web at Scale: 3 Months of Failures — oatmealdaddy4 · 2026-10-04
- Open-source MCP memory server CRBRO passes 12/12 real-agent cross-session tests — AntonioJBer · 2026-10-04
- Simon Willison: We Need Default Hard Budget Caps on AI Services — elffjs · 2026-10-04