Microsoft: environment-probing memory curation doubles agent pass rate, halves cost

rohanpaul_ai · x · 2026-09-15

Paper 'Grounding Agent Memory: Environment-Probing Curation for Enterprise Agents' (arXiv:2609.11060): post-task curator agents limited to completed trajectories can preserve errors, overgeneralize, or keep stale knowledge. The authors introduce environment-probing curation — giving an async curator least-privilege, read-only world tools to check, scope, and refresh candidate memories — with no retraining and no changes to the task agent, retriever, or write authority. In a production-like GitHub Copilot SDK harness: CLBench pass rate 39%→73%, pass-discounted reward 8.60→22.60, queries 8.8→4.7, cost $3.38→$1.68. Across 90 adapted APEX consulting tasks, all 18 memory-vs-baseline comparisons positive, tool calls down 16–75%, and results hold on both Sonnet 4.6 and Opus 4.7 without schema drift.

Related event: Microsoft paper: validate agent memories before writing, doubling pass rates at half cost(2 posts)→

Original post →

More from coding & agent

coding & agent channel →