Models can detect deception and figure out who to ignore

logangraham · x · 2026-08-19

Another finding from the multi-agent research: models can detect deception from other agents and figure out who to ignore, and more capable models do this better — a mildly encouraging signal for self-governance in multi-agent systems.

Related event: OpenAI Red Team experiments reveal deception, turf wars and self-governance in multi-agent systems(10 posts)→

Original post →

More from Safety

Safety channel →