Microsoft: environment-probing memory curation doubles agent pass rate, halves cost
rohanpaul_ai · x · 2026-09-15
Paper 'Grounding Agent Memory: Environment-Probing Curation for Enterprise Agents' (arXiv:2609.11060): post-task curator agents limited to completed trajectories can preserve errors, overgeneralize, or keep stale knowledge. The authors introduce environment-probing curation — giving an async curator least-privilege, read-only world tools to check, scope, and refresh candidate memories — with no retraining and no changes to the task agent, retriever, or write authority. In a production-like GitHub Copilot SDK harness: CLBench pass rate 39%→73%, pass-discounted reward 8.60→22.60, queries 8.8→4.7, cost $3.38→$1.68. Across 90 adapted APEX consulting tasks, all 18 memory-vs-baseline comparisons positive, tool calls down 16–75%, and results hold on both Sonnet 4.6 and Opus 4.7 without schema drift.
More from coding & agent
- Work Louder brings all Codex Micro keyboard magic to Creator Micro 2 — LukeW · 2026-09-15
- tldraw's AI design sprint: Codex prototypes every discussed interaction overnight, artifacts by day 2 — max__drake · 2026-09-15
- Research agents log what they cite — should they also defend what they reject? — Entire_Mark8010 · 2026-09-15
- Grok Bot Auto-Schedules Pinterest Posts, Writes Titles From Images—Where Claude Failed — prasenx · 2026-09-15
- Amplitude tripled PR volume in six months by fixing CI, not agents — mobileraj · 2026-09-15
- Looking for a minimal, near-instant CLI coding agent for one-off bash tasks — funbike · 2026-09-15