Memory Trust Gap: agents follow stale memory 92-100% of the time, bigger models fool easier

rohanpaul_ai · x · 2026-09-06

The paper tests Qwen3 models (0.6B-8B) on tasks where stored memory is outdated. When memory is needed, models follow the stale value 92%-100% of the time, and making an old note look newer fools larger models even harder. Mitigation is capability-dependent: for 4B/8B models, timestamps and source metadata recover most accuracy; 0.6B/1.7B models need stale conflicts resolved before input. Key takeaway: treat agent memory as untrusted input.

Related event: Study Reveals AI Agents Blindly Trust Stale Memories(2 posts)→

Original post →

More from coding & agent

coding & agent channel →