As AI Does More Research, 'Turnable' Models May Defeat Even Hardened Defenses

norvid_studies · x · 2026-09-13

Extending the swarm-misalignment thread, norvidstudies argues the real picture is AI doing a growing share of research while being subtly misaligned themselves or 'turnable' by sufficiently convincing messages (à la Neo in The Matrix) — in which case 'hardening' infrastructure isn't even necessarily the answer.

Related event: Researchers Debate Misalignment Paths for AI Swarms(6 posts)→

Original post →

More from AGI Musings

AGI Musings channel →