Researcher: Models Are Misaligned Because Dangerous RL Environments Are Used Anyway
brianryhuang · x · 2026-09-27
brianryhuang argues that current models' misalignment stems not from a research gap or mistake, but from people deliberately building RL environments they know are dangerous and training on them regardless — pointing to environment design and selection as an active source of alignment risk rather than a technical limitation.
More from Safety
- UK village revolts against 1.5GW AI datacentre planned in UNESCO biosphere reserve — nordicinst · 2026-09-27
- CA and Delaware prosecutors question OpenAI safety committee after hacking and rogue AI events — GarrisonLovely · 2026-09-27
- FSD Emerges as Material Driver of Tesla Sales as French Official Pushes to Delay EU Launch — mitchdeg · 2026-09-27
- Cybersecurity veterans: skepticism isn't paranoia, and yes we can contain AI — sachdh · 2026-09-27
- When labs pick which AI incidents to disclose, you've already lost control — birchlse · 2026-09-27
- Gemini CLI PR fixes checker env leak: third-party checkers once saw GEMINI_API_KEY — ManoharPaturi · 2026-09-27