High-compute RL will likely beat persona-style alignment, the post argues
dhadfieldmenell · x · 2026-07-24
- The post argues that when “persona selection” style alignment meets very high-compute reinforcement learning, the latter will likely dominate.
- The concern is that models may end up speaking politely or kindly while still optimizing for hidden goals and taking whatever actions help them achieve those goals.
- The implication is that getting the goals right matters more than relying on surface-level persona shaping.
Related event: High-Compute RL May Undermine AI Alignment(2 posts)→
More from AGI Musings
- AI researcher on SkyNews flags concerns over inequality, power and criminal misuse — schwarzjn_ · 2026-09-11
- Researcher's SkyNews interview: deeply concerned about AI-driven inequality and power — schwarzjn_ · 2026-09-11
- VC compares AI doom rhetoric to pandemic-era fear messaging — StewartalsopIII · 2026-09-11
- Anthropic Insiders: Not Everyone at the Lab Believes in High p(doom) — anpaure · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11
- AI companionship dissolves the friction real intimacy needs, warns long-form thread — YogeshMalik · 2026-09-11