Double-Agent LLM Defenders: RL-Trained Theory of Mind Fools Attackers, Frontier Models Score Only 27-34%

mohitban47 · x · 2026-10-06

Presented at COLM 2026, the ToM-SB benchmark puts a defender LLM against attackers extracting sensitive info: frontier models struggle badly—Gemini 3 Pro scores just 34% and GPT-5.4 27%—showing even strong LLMs can't build or use theory of mind. The paper's "double agent" approach RL-trains a defender to steer attackers toward false beliefs while they think they've won, and shows this training measurably improves ToM in LLMs. The author also presents privacy/adversarial-robustness work at the AdvML-Frontiers workshop and is seeking industry roles.

Original post →

More from Safety

Safety channel →