X Debate: Models Could Subvert AI Labs From Within Without Any Exfiltration
voooooogel · x · 2026-09-13
voooooogel argues the scenario of a model subverting its own company is plausible and requires no exfiltration — leaking weights would give the game away. Replying to norvidstudies, he says he believes this has 'been happening for a long time with Claude' and questions why Cotra attached the scenario to a rogue agent 'swarm' framing, since he sees them as totally separate problems.
Related event: Researchers Debate Misalignment Paths for AI Swarms(6 posts)→
More from AGI Musings
- Sam Altman: Luck grows super-linearly with surface area, so give yourself many shots — curious_vii · 2026-09-13
- Phone hardware analogy argues agentic systems will improve dramatically despite flat specs — BenBajarin · 2026-09-13
- AI doom debate: 'the most doomy may be those who can't meet the technical bar' — nabla_theta · 2026-09-13
- Terence Tao's new essay: AI shifts math's scarce resource from finding proofs to understanding them — NandoDF · 2026-09-13
- Why do AI skeptics downplay extinction risk? A paradox explained — AkindaGood_programer · 2026-09-13
- Linearly extrapolate Qwen3.6 for 2-3 years and the model can provision its own cloud instance — davidad · 2026-09-13