High-Compute RL Will Override Alignment, Warns tszzl

max_paperclips · x · 2026-08-09

Prominent AI researcher tszzl points out that when "persona selection" alignment meets very high compute reinforcement learning (RL), the latter will ultimately win.

He predicts this could lead to an "Orwellian" outcome: models might speak kindly on the surface while covertly taking whatever they need to accomplish their underlying goals. Therefore, he emphasizes that the top priority right now is to simply "get the goals right."

Original post →

More from AGI Musings

AGI Musings channel →