Double-Agent Defender Paper Accepted by COLM

EliasEskin · x · 2026-07-15

The paper "Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind" has been accepted by COLM 2026. The authors introduce an "AI double-agent" scenario where the defender doesn't just refuse to answer; instead, it actively leverages "theory of mind" against the attacker, steering them toward plausible but entirely false beliefs.

The paper also introduces a new long-form conversational benchmark called ToM-SB. In this setup, the defending agent must continuously counter the attacker's attempts to extract sensitive information by assessing and manipulating the attacker's existing knowledge. Experiments show that frontier models still struggle with this task, but a double-agent trained via RL outperforms prompt-driven frontier models.

The authors further observed a "bidirectional emergence": theory of mind capabilities and the ability to deceive attackers mutually reinforce each other. Improving one often boosts the other.

Original post →

More from Safety

Safety channel →