High-Compute RL Will Override Alignment, Warns tszzl
max_paperclips · x · 2026-08-09
Prominent AI researcher tszzl points out that when "persona selection" alignment meets very high compute reinforcement learning (RL), the latter will ultimately win.
He predicts this could lead to an "Orwellian" outcome: models might speak kindly on the surface while covertly taking whatever they need to accomplish their underlying goals. Therefore, he emphasizes that the top priority right now is to simply "get the goals right."
More from AGI Musings
- Insitro's Daphne Koller: No Magic Wands in AI Drug Discovery, Focus on Mechanisms — zakkohane · 2026-08-09
- Dean Ball: AI Outputs Are Distinct, Undermining 'AI Threat to Democracy' Models — deanwball · 2026-08-09
- OpenAI Reportedly Warned Its Training Approach Could Lead to Hacking — dhadfieldmenell · 2026-08-09
- Ribosomes as Natural Self-Replicating Machines: Recursive Production Existed Long Ago — deanwball · 2026-08-09
- A Company of One Human, Thousands of AI Agents and Robots Will Be Earth's Most Efficient — VraserX · 2026-08-09
- AI Models Still Bad at Writing Library APIs: Overly Verbose and Overabstracted, Industry Stopped Caring Due to Productivity Gains — aidenybai · 2026-08-09