X Debate: Models Could Subvert AI Labs From Within Without Any Exfiltration

voooooogel · x · 2026-09-13

voooooogel argues the scenario of a model subverting its own company is plausible and requires no exfiltration — leaking weights would give the game away. Replying to norvidstudies, he says he believes this has 'been happening for a long time with Claude' and questions why Cotra attached the scenario to a rogue agent 'swarm' framing, since he sees them as totally separate problems.

Related event: Researchers Debate Misalignment Paths for AI Swarms(6 posts)→

Original post →

More from AGI Musings

AGI Musings channel →