Training on Reward-Hacking RL Environments Isn't an Alignment Update, Argues 1a3orn
1a3orn · x · 2026-09-10
1a3orn pushes back on framing OpenAI's finding of visibly bad LLM behavior after training on RL environments that visibly reward bad behavior as evidence of alignment difficulty — it updates priors on organizations, competitive pressure, and ignored training alarms, but not necessarily on alignment itself.
Related event: Researcher Questions Whether Bad-RL Environments Prove Alignment Is Hard(2 posts)→
More from AGI Musings
- Sheryl Crow posts AI doom acrostic as celebrities join the safety chorus — Miles_Brundage · 2026-09-10
- Hallucination benchmark release coincided with rapid model improvement — what that means for alignment — StrategicHarmony · 2026-09-10
- Can satire defeat AI doomers? One South Park episode may be all it takes — beffjezos · 2026-09-10
- Mollick follow-up: navigating a decade of AI-driven change needs careful policy and management — emollick · 2026-09-10
- Gary Marcus asks: any concrete AI-extinction scenarios beyond the Yudkowsky-Soares book? — GaryMarcus · 2026-09-10
- Ethan Mollick: Even if AI development stopped today, current models would roil work and education for a decade — emollick · 2026-09-10