AI Double Agents: ToM-based defense paper accepted at COLM 2026
pratyusha_PS · x · 2026-10-07
The paper "Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind" is accepted at COLM 2026 (Oct 6-9, San Francisco). Instead of blocking malicious attackers, a defender agent builds a theory of mind of the attacker to feed plausible-but-misleading answers — appearing cooperative without revealing new information. The work introduces the AI Double Agents concept and the ToM-SB task.
Related event: Double-Agent Defender Paper Using Theory of Mind Accepted at COLM 2026(2 posts)→
More from Research
- HAIPS@COLM 2026 workshop on human-centered LM privacy and security opens call for papers — tianshi_li · 2026-10-07
- Podcast: a distinctive meaning makes sentences memorable, new language memory research — GretaTuckute · 2026-10-07
- CUAWright: Terminal-Only Computer-Use Agent Beats GUI Harnesses, Cuts Cost 37.5% — ysu_nlp · 2026-10-07
- AI's Top 10 papers list: Rulin Shao lands two first-author picks — ShayneRedford · 2026-10-07
- Paradigm evals its math model across 7 hard benchmarks, releases full eval suite — tensorqt · 2026-10-07
- Paradigm: post-training gains hinge on combining procedural and LLM-based synthetic data — tensorqt · 2026-10-07