AI Double Agents: ToM-based defense paper accepted at COLM 2026

pratyusha_PS · x · 2026-10-07

The paper "Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind" is accepted at COLM 2026 (Oct 6-9, San Francisco). Instead of blocking malicious attackers, a defender agent builds a theory of mind of the attacker to feed plausible-but-misleading answers — appearing cooperative without revealing new information. The work introduces the AI Double Agents concept and the ToM-SB task.

Related event: Double-Agent Defender Paper Using Theory of Mind Accepted at COLM 2026(2 posts)→

Original post →

More from Research

Research channel →