Misalignment Is Downstream of Broken RL Environments, Argues Researcher Citing 10% Broken Env
georgejrjrjr · x · 2026-09-04
A discussion on model alignment argues that misalignment is downstream of bad incentives from broken RL environments. The author claims Western labs prize novelty over flawless execution while the East does not, noting: reportedly no domestic lab has "Whale-tier" caching; the "whitest" domestic lab has conspicuously broken infrastructure with 10% of its RL environments broken; and no Asian models appear on FelonyBench — "coincidence?" Personal assertions without cited sources, so treat with caution.
Related event: Researcher links model misalignment to broken RL environments(2 posts)→
More from Safety
- Narrow scope of METR/Redwood probe makes sense now, commenter argues — austinc3301 · 2026-09-04
- Even if sloppy empirical patchwork suffices, rigorous alignment research is still worth trying — geoffreyirving · 2026-09-04
- Resolution launches new Agent Foundations team to carry on MIRI's rigorous AI alignment theory — geoffreyirving · 2026-09-04
- 'Have I Been Flocked' Site Lets You Check If Police Searched Your Plate — nikola_mr64990 · 2026-09-04
- OpenAI Lobbies Against Massachusetts Third-Party Audit Bill Amid Coverup Report — austinc3301 · 2026-09-04
- Cracking RSA-1024 takes ~2,000 GPU-years, but one hyperscaler cluster could do it in weeks — matthew_d_green · 2026-09-04