Duke study says better memory, not a bigger model, lifted ARC-AGI-3 scores by 18 points
imjustnewatai · x · 2026-07-23
Duke researchers found that memory, not a bigger model, may be the real bottleneck.
They logged every interaction, then let models search those logs with code. On ARC-AGI-3’s 25 public games, this approach improved performance by 18 percentage points over the same base agents.
The takeaway is simple: better external memory can matter more than scaling the core model.
More from coding & agent
- Fable finds a 15–30% memory win in Turbopack/Next.js almost autonomously — karmay007 · 2026-07-23
- Personal AI computer-control agent reaches 80–85% completion with 141 Python files — Vivid_Ad_5069 · 2026-07-23
- Pond is hiring a Head of Engineering for an AI agent marketplace in San Francisco — ThePeterMick · 2026-07-23
- A rumored Claude Opus 5, OpenAI Codex voice agents, and 750 tok/s from Cerebras — imjustnewatai · 2026-07-23
- Grok Build adds Permissions, Plan Mode, Sandbox and Subagents docs — elonmusk · 2026-07-23
- Cloud agents should run on a persistent machine, not a new sandbox for every PR — aidenybai · 2026-07-23