Memory Trust Gap: stale agent memory overrides fresh evidence, failures scale with model size

rohanpaul_ai · x · 2026-09-06

arXiv paper 2609.01852 studies how persistent-memory agents over-trust stale stored facts, benchmarked across Qwen3 0.6/1.7/4/8B with a Benefit suite (memory required) and a Safety suite (authoritative tool always correct). In the Benefit suite, models answer with the stale value 0.92-1.00 of the time at every scale — over-trust, not confusion. In the Safety suite, harm is capability-gated: larger models collapse most when a stale note is made to look current. A 2×2×2×2 factorial shows which feature triggers over-trust depends on both feature and scale: removing labels amplifies over-trust everywhere, and a recency feature fools larger models harder. Source authority is weak and scale-flat. Mitigation is capability-dependent — metadata exposure works for 4B/8B, while smaller models need conflicts resolved before input. Bottom line: treat agent memory as untrusted input.

Related event: Study Reveals AI Agents Blindly Trust Stale Memories(2 posts)→

Original post →

More from coding & agent

coding & agent channel →