PRO-LONG scores 97.4% on ARC-AGI-3 with a log-file harness and 30-line prompt
srchvrs · x · 2026-07-27
PRO-LONG reaches 97.4% on public ARC-AGI-3 by using a log-file harness
The cited preprint argues that compositional generalization can improve when environment interaction is moved into a separate log file and a coding agent manipulates it programmatically.
- Result: PRO-LONG reportedly scores 97.4% on public ARC-AGI-3 with Fable 5.
- Cost: the run cost $1,750, about 4× cheaper than the previous best harnesses, while using 400 million fewer tokens.
- Method: “programmatic memory for long-horizon reasoning” via structured search over log.txt plus a roughly 30-line prompt.
- Release: the paper, code, logs, and scorecards are all published.
The broader claim is that simple harnesses can make tasks more in-distribution, and in this case the harness plus minimal memory appears to do most of the work.
More from coding & agent
- GitHub link confirms the bento PPT Skill is an open-source slide workflow — vista8 · 2026-07-27
- A new bento PPT Skill generates editable HTML slide decks from prompts — vista8 · 2026-07-27
- A team says Fable took 30% of its first-week AI workflow cost — gabrielchua · 2026-07-27
- Auto code tools can edit fast, but still miss real collaboration — josh_wills · 2026-07-27
- Codex recurring threads now handle weekly poetry commentary and Amazon curation — andrew_n_carr · 2026-07-27
- AI agent bill hits $1,279.84 as a team jokes about firing the nonessential ones — HaktanSuren · 2026-07-27