Double-Agent LLM Defenders: RL-Trained Theory of Mind Fools Attackers, Frontier Models Score Only 27-34%
mohitban47 · x · 2026-10-06
Presented at COLM 2026, the ToM-SB benchmark puts a defender LLM against attackers extracting sensitive info: frontier models struggle badly—Gemini 3 Pro scores just 34% and GPT-5.4 27%—showing even strong LLMs can't build or use theory of mind. The paper's "double agent" approach RL-trains a defender to steer attackers toward false beliefs while they think they've won, and shows this training measurably improves ToM in LLMs. The author also presents privacy/adversarial-robustness work at the AdvML-Frontiers workshop and is seeking industry roles.
More from Safety
- McDonald's hit with class action for allegedly using AI and nonpublic data to coordinate menu prices across ~14,000 US restaurants — Polymarket · 2026-10-06
- AI expert on ABC: policy isn't innovation vs safety, voluntary safeguards need teeth — chrismattmann · 2026-10-06
- Oxford OII maps three risks of AI agents in government and five principles to keep humans in charge — _akpiper · 2026-10-06
- White House launches Genesis Mission, billed as an AI Manhattan Project — aKaizuh · 2026-10-06
- MCP agent-to-agent comms may be the riskiest protocol you've never heard of — Ars Technica AI · 2026-10-06
- DEF Report: The Shared Developer Liability Question for AI and Crypto Code — Prof_CarlaReyes · 2026-10-06