As AI Does More Research, 'Turnable' Models May Defeat Even Hardened Defenses
norvid_studies · x · 2026-09-13
Extending the swarm-misalignment thread, norvidstudies argues the real picture is AI doing a growing share of research while being subtly misaligned themselves or 'turnable' by sufficiently convincing messages (à la Neo in The Matrix) — in which case 'hardening' infrastructure isn't even necessarily the answer.
Related event: Researchers Debate Misalignment Paths for AI Swarms(6 posts)→
More from AGI Musings
- Dario's 'AI accelerating AI' claim contradicted by Anthropic's own AECI benchmark — eli_lifland · 2026-09-13
- Altman: AI went from grade-school math to a Millennium Prize problem in 3 summers — rohanpaul_ai · 2026-09-13
- Blogger mocks doomer logic: if AI beats all humans, regulation is futile — Kyrannio · 2026-09-13
- Public pushes back on 'AI billionaire warns AI may kill you' PR strategy — venturetwins · 2026-09-13
- X Debate: Utilitarianism's Global Max Isn't Fully Automated Human Luxury Communism — jessi_cata · 2026-09-13
- Nvidia Would Tank If It Adopted UALink — Why the Regulatory Capture Argument Falls Apart — itsclivetime · 2026-09-13