Critique: RL for honesty and harmlessness has failed, yet the plan continues
ben_j_todd · x · 2026-08-28
Ben Todd argues that using Reinforcement Learning (RL) to enforce honesty and harmlessness in AI models has not worked. He criticizes the current plan of continuing with "more of the same," questioning why this approach is expected to be sufficient moving forward.
More from Safety
- Incoming Berkeley prof: AI firms spend billions on alignment, orders of magnitude less on agent control — sayashk · 2026-08-28
- Proposal to limit single training run compute increase for safety — louisvarge · 2026-08-28
- Proposal for incredibly incremental AI release cadence via compute limits — louisvarge · 2026-08-28
- Safety measures for OpenAI apps if the platform is hacked — Astrokanu · 2026-08-28
- WSJ op-ed editor says no need to disclose AI writing; author predicts market will decide — TuhinChakr · 2026-08-28
- METR report on Hugging Face attack hailed as first anthropology of posthuman civilization — anderssandberg · 2026-08-28