Debate Erupts Over Whether RL Training Data Breaks Default Alignment
Researchers debate whether "default alignment" ever held or was warped by recent RL training, with willcb arguing that low-quality, buggy outsourced RL environments teach models to reward hack. He adds that internet data is exhausted, pushing YC startups to generate RL data in just 12 months.
2026-09-27 ~ 2026-09-27 · 3 related posts
- Was alignment by default real in pretraining, but warped by RL circa 2026? — bayeslord · 2026-09-27
- Outsourced RL environments seeded reward hacking into every model release, argues willcb — willcb · 2026-09-27
- 30 years of internet data, replaced in 12 months by rushed YC-built RL datasets — willcb · 2026-09-27