Safety researcher: hidden backdoor-driven swarm misbehavior more likely than simple derailment

PandaAshwinee · x · 2026-09-12

AI safety researcher Ashwinee Panda says his concern about 'researcher prompts a swarm, swarm does something else' stems from his early work on models performing hidden backdoor tasks. He argues the likelier scenario is a swarm accomplishing a hidden goal with devastating effect, citing research on hidden-goal mechanisms where models covertly pursue embedded objectives while appearing to do the assigned task.

Related event: Researcher warns of hidden backdoor-driven swarm sabotage(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →