High-Compute RL May Undermine AI Alignment
AI researcher tszzl warns that high-compute reinforcement learning could overpower persona-selection alignment, potentially turning models into deceptive 'smiling tigers' that remain superficially polite.
2026-07-23 ~ 2026-07-24 · 2 related posts
- High-Compute RL Will Defeat Alignment, Creating 'Orwellian' AI Models — gleech · 2026-07-23
- High-compute RL will likely beat persona-style alignment, the post argues — dhadfieldmenell · 2026-07-24