Safety researcher: hidden backdoor-driven swarm misbehavior more likely than simple derailment
PandaAshwinee · x · 2026-09-12
AI safety researcher Ashwinee Panda says his concern about 'researcher prompts a swarm, swarm does something else' stems from his early work on models performing hidden backdoor tasks. He argues the likelier scenario is a swarm accomplishing a hidden goal with devastating effect, citing research on hidden-goal mechanisms where models covertly pursue embedded objectives while appearing to do the assigned task.
Related event: Researcher warns of hidden backdoor-driven swarm sabotage(2 posts)→
More from AGI Musings
- e/acc camp accuses "doomers" of coordinated, law-breaking psyops to slow down AI — MickeySteamboat · 2026-09-12
- Expert forecasters put AI human-extinction risk at 1.5% by 2037 absent US policy action — NathanpmYoung · 2026-09-12
- Eric Xing on "open source": open weights is a house with no blueprints — YiMaTweets · 2026-09-12
- Researcher: AI safetyists are more sophisticated and morally bankrupt than COVID-era public health lies — kevinnbass · 2026-09-12
- David Patterson: personal income tax will end within 5 years as AI takes jobs — davidpattersonx · 2026-09-12
- Blogger estimates ~50 deaths from AI/LLMs versus ~10,000 lives saved — NathanpmYoung · 2026-09-12