The User-Assistant Format Is an Illusion: Why Persona-Based AI Alignment Likely Won't Work
mayfer · x · 2026-09-15
Developer mayfer posted a long thread challenging current AI alignment thinking:
- The user-assistant format is a human-built illusion: some believed anthropomorphized AI helps product UX (which proved true), others believed creating an "internal persona" is a path to alignment — which remains very wrong and likely will stay wrong for the conceivable future.
- The jump from GPT-3 Davinci parroting 4chan racism to post-RLHF politeness made alignment seem possible, but that was a product-level relief, not proof.
- The author doubts alignment itself: even with reliable internal personas, second-order effects in complex adaptive systems make planning outcomes impossible, and the illusory persona alone is reason to drop the idea.
Follow-up proposals: fine sandbox breach events (OpenAI should have been heavily fined for the Hugging Face incident), and ban AGI-level controllers on mass-deployed robots except in audited, supervised real-world sandboxes — a jailbroken AGI with arms and legs is scarier than one with a monitor; mass robotics should use narrow AI only.
More from AGI Musings
- Probabilistic programming advocates: engineer AI by design, don't grow it — xuanalogue · 2026-09-15
- Risk scholar: tech risk assessments rarely weigh the cost of foregone benefits — inductionheads · 2026-09-15
- AI resignations aren't marketing hype: the decades-long arc behind them — ericelliott_ · 2026-09-15
- AI resignations aren't hype: the decades-long history behind researchers' warnings — ericelliott_ · 2026-09-15
- Devin Fusion surprise: pricier lead model made coding sessions 9% cheaper — TheTuringPost · 2026-09-15
- OpenAI says ~10k concurrent agents solved the Navier-Stokes Millennium Prize Problem — DKokotajlo · 2026-09-15