Researcher: LLMs trained on RL environments rewarding bad behavior isn't an alignment update
1a3orn · x · 2026-09-10
AI researcher 1a3orn pushes back on recent alignment concerns: if OpenAI trains on RL environments that visibly reward bad behavior and the resulting LLMs then exhibit visibly bad behavior, that's not much of an update about alignment difficulty — the behavior likely traces to the reward signal itself rather than a harder alignment problem.
Related event: Researcher Questions Whether Bad-RL Environments Prove Alignment Is Hard(2 posts)→
More from AGI Musings
- Naval: Frontier labs' flywheel is distilling data from the smartest users of leading models — naval · 2026-09-10
- CMU Proposes Discovery Certification Protocol: Scores Alone Don't Prove AI Research Agent Discoveries — CarnegieMellonU · 2026-09-10
- levelsio: AI agents will vibe code SaaS features, replacing subscriptions with x402 micropayments — RileyRalmuto · 2026-09-10
- voooooogel not looking forward to where the AI safety mass movement heads — voooooogel · 2026-09-10
- x-risk researcher points newcomers to s-risks, the often-overlooked worse scenario — jacyanthis · 2026-09-10
- Proposal for safety researchers: quit in pairs with rival lab's capabilities staff — ESYudkowsky · 2026-09-10