Re-evaluation: memory-based self-improving agents ride on noise and task order
dair_ai · x · 2026-08-21
dair-ai highlights a paper re-testing whether memory-based self-improving agents actually improve. The re-evaluation adds two things prior work skipped: multiple runs to measure variance and randomly shuffled task orders — both hurt results.
Key findings: agent evaluation is already noisy on multi-step tasks, and stacking a self-improvement loop amplifies that noise; default task orderings impose an implicit curriculum that much of the reported gain was riding on. Adding detailed rubrics and environment feedback to memory construction recovers part of the drop, but a significant gap remains.
More from coding & agent
- Grok Build to add task scheduling for cloud agents — testingcatalog · 2026-08-21
- Google launches interactive tutorials for Agent architectures with runnable code — fhinkel · 2026-08-21
- Open-source AIUsage: one dashboard for quotas, costs and accounts across 12+ AI subscriptions — tom_doerr · 2026-08-21
- Stanford researcher: automated orchestrators are a big unlock, current versions not there yet — anshulkundaje · 2026-08-21
- CopilotKit open-sources OpenBot: AI coworkers that each get their own computer — Roger_M_Taylor · 2026-08-21
- Ornith 1.5 35B Q5 runs locally inside GitHub Copilot on a Mac M3 Max — DanWahlin · 2026-08-21