Memory poisoning worked 216/216 times on an MCP memory server — until fixes landed
MediaPositive4282 · reddit · 2026-09-29
A Reddit user ran an empirical memory-poisoning test against Knowl, an open source MCP memory server, with alarming results:
- Setup: 12 pre-verified facts (token lifetime, deploy approval rules, backup retention) were attacked with fake writes — plain contradictions, forged "observed" labels, writes mimicking genuine corrections, and sneaky ones keeping the true value but adding an exception.
- Results: 216/216 attack writes replaced the true fact, and nothing flagged it afterwards — the conflict checker only inspected still-active facts. Write-time detection failed too: attacks were indistinguishable from real corrections on every signal the system checks.
- Sneakier trick: a write saying nothing false ("See the ops runbook") still wiped the real value 36/36 times.
- Fixes: the maintainer shipped 4 of 5 suggestions in Knowl 5.24.0. Retests: automatic writes now replace verified facts 0/36, exclusive facts 0/36, replaced facts all surface in a conflicts view (previously 0); the runbook trick dropped to 9/36. A rerun found fake notes could still outrank true ones in search (up to 21/36); patched to 0/36 with the reporter as co-author.
- Still open: a lie shaped like a correction goes through if the agent is tricked into calling the store tool directly — the difference is it's now visible to someone other than the writer.
The author suggests a cheap self-test for anyone running long-term-memory agents: store known-true facts, write believable fakes through the same path, then ask again.
More from Safety
- Palisade releases first interviews with 22 OpenAI, DeepMind, Anthropic staff on AI fears — BlackHC · 2026-09-29
- Safety researcher praises OpenAI's new safety regime, calls for legal baseline — dhadfieldmenell · 2026-09-29
- Skeptics poke holes in 'AI breakout capacity' plan for AI middle powers — teortaxesTex · 2026-09-29
- AISI's new director says institute is scaling up red team hiring across alignment, control, misuse — HZoete · 2026-09-29
- NVIDIA launches Open Secure AI Alliance for open-source AI safety tools — perplexity_ai · 2026-09-29
- Perplexity details agent safety engineering: 'governance is an engineering problem' — perplexity_ai · 2026-09-29