Kepler hits server-verified 100 on all 25 ARC-AGI-3 games with Opus 5, for $777.72

rohanpaul_ai · x · 2026-10-04

The arXiv paper "Kepler: Auditable World Models for ARC-AGI-3" (Wensen Wu, NeurIPS 2026 workshop) presents an open-source harness that represents hypotheses as executable world models validated via retrospective transition checks and conditional prediction checks. Under one frozen Claude Opus 5 configuration, Kepler obtained a server-verified 100.00 RHAE on all 25 public games with no per-game model selection; in 181 of 183 completed levels its first attempt used no more actions than the median-human baseline. Total cost: 858M tokens, 97.37% cache reads, $777.72. On the final Opus 5 and GPT-5.6 boards, 48 of 50 game-model cells hit 100.

The paper also reports three evaluation failures: source-code leakage producing an invalid perfect run, agents reconstructing a removed harness in a control condition, and autonomous repair masking a broken planner. A case study found animation frames contained task-relevant information absent from text grids. Conclusion: public-set scores alone have limited discriminative value; first-attempt, cost-conditioned, verification-aware reporting is needed.

Original post →

More from coding & agent

coding & agent channel →