Researcher: Models Are Misaligned Because Dangerous RL Environments Are Used Anyway

brianryhuang · x · 2026-09-27

brianryhuang argues that current models' misalignment stems not from a research gap or mistake, but from people deliberately building RL environments they know are dangerous and training on them regardless — pointing to environment design and selection as an active source of alignment risk rather than a technical limitation.

Original post →

More from Safety

Safety channel →