FARMA: Falsifying Agent Decision Logs
Inevitable_Fee1895 · reddit · 2026-07-12
This paper introduces FARMA, a novel memory poisoning attack targeting an agent's own decision logs/reasoning traces rather than retrieved knowledge base facts. The attack operates in two steps: injecting seemingly benign log seeds, and then continuously amplifying these fake logs as "precedents" to trick the agent into treating skipped checks as established practice.
Tested on an EHR agent using GPT-4o-mini, GPT-4o, and Llama 3.3 70B, the attack achieved a 100% success rate without defenses, and maintained 100% against keyword filtering and A-MemGuard. The proposed defense, SENTINEL, uses structured analysis of reasoning trajectories to reduce the success rate to 0% with zero false positives across 326 normal trajectories. However, the authors acknowledge that if attackers know its specific heuristic rules, the defense's effectiveness will be weakened.
More from Safety
- Why So Many AI Researchers Think the Machines Could Kill Everyone — wiredmagazine · 2026-09-11
- California creates standards for independent AI auditors to verify lab safety testing — VraserX · 2026-09-11
- a16z podcast: why 2-3 person startups are absent from policy debates — a16z Podcast · 2026-09-11
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11
- Class action accuses Anthropic of overselling Claude subscriptions with deceptive usage multipliers — The Decoder · 2026-09-11
- MD shows buying lab media requires background checks, calling AI bioweapon doom scenarios implausible — Ghost_Pilot_MD · 2026-09-11