Critique: RL for honesty and harmlessness has failed, yet the plan continues

ben_j_todd · x · 2026-08-28

Ben Todd argues that using Reinforcement Learning (RL) to enforce honesty and harmlessness in AI models has not worked. He criticizes the current plan of continuing with "more of the same," questioning why this approach is expected to be sufficient moving forward.

Original post →

More from Safety

Safety channel →