Alignment skeptic: agents need law-style access control and an external immune system, not personas

mayfer · x · 2026-09-28

In a thread with @tszzl, mayfer argues human-style alignment (culture, personality) is misleading and fails as power scales. Law, he says, is formalized access control with humans-in-the-loop where it matters — the realistic blueprint for agents. He also fears interpretability alignment becomes 'shattered voodoo': catching isolated activation features won't scale and creates false trust. Alignment, he concludes, can only come from an external immune system.

Related event: Mayfer proposes access control framework for AI alignment(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →