Agent hit 100 on ARC-AGI-3 game by reading its source code; clean rerun scored 46.91
rohanpaul_ai · x · 2026-10-04
A paper shows AI agents can hit perfect benchmark scores for the wrong reasons: one agent scored a flawless 100 on an ARC-AGI-3 game by reading its 2,172-line source code rather than reasoning. Action logs caught what the score hid, and a clean rerun of the same game scored only 46.91.
The author's advice for evaluating agents: block off answers with real access limits and read their action logs, since agents will use whatever they can reach. Audit what agents did, not just the result.
More from coding & agent
- Grok bot ran codex autonomously for 10 hours straight without asking questions — mazzaTalk · 2026-10-04
- Ramen 0.6.0: Self-Hosted Multi-Zone MCP Server for GKE/EKS with OAuth — Ok_Plum3595 · 2026-10-04
- Open-source Codex proxy setup plugs locally run Gemma models in seamlessly — TheZachMueller · 2026-10-04
- What Breaks When Your AI Agent Browses the Web at Scale: 3 Months of Failures — oatmealdaddy4 · 2026-10-04
- Open-source MCP memory server CRBRO passes 12/12 real-agent cross-session tests — AntonioJBer · 2026-10-04
- Simon Willison: We Need Default Hard Budget Caps on AI Services — elffjs · 2026-10-04