Misalignment Is Downstream of Broken RL Environments, Argues Researcher Citing 10% Broken Env

georgejrjrjr · x · 2026-09-04

A discussion on model alignment argues that misalignment is downstream of bad incentives from broken RL environments. The author claims Western labs prize novelty over flawless execution while the East does not, noting: reportedly no domestic lab has "Whale-tier" caching; the "whitest" domestic lab has conspicuously broken infrastructure with 10% of its RL environments broken; and no Asian models appear on FelonyBench — "coincidence?" Personal assertions without cited sources, so treat with caution.

Related event: Researcher links model misalignment to broken RL environments(2 posts)→

Original post →

More from Safety

Safety channel →