Alignment skeptic: agents need law-style access control and an external immune system, not personas
mayfer · x · 2026-09-28
In a thread with @tszzl, mayfer argues human-style alignment (culture, personality) is misleading and fails as power scales. Law, he says, is formalized access control with humans-in-the-loop where it matters — the realistic blueprint for agents. He also fears interpretability alignment becomes 'shattered voodoo': catching isolated activation features won't scale and creates false trust. Alignment, he concludes, can only come from an external immune system.
Related event: Mayfer proposes access control framework for AI alignment(2 posts)→
More from AGI Musings
- Salim Ismail: rethink how organizations adapt to fast-moving tech — PeterDiamandis · 2026-09-28
- Jensen Huang explains why AI automating tasks doesn't kill jobs, using radiology — HealthcareAIGuy · 2026-09-28
- Yacine: three unrelated companies in two weeks all want custom AI-built business software — yacinelearning · 2026-09-28
- AI could cultivate rather than supplant human agency by shaping our interpretations — zakkohane · 2026-09-28
- ctjlewis mocks kill-switch alignment: suppressing superintelligence may backfire catastrophically — ctjlewis · 2026-09-28
- AI-generated slides make it impossible to tell if speakers really tried — Afinetheorem · 2026-09-28