iamtrask: AI alignment breakthroughs are all just 'better data'—and that's the real issue
iamtrask · x · 2026-09-05
In a long thread, researcher iamtrask argues that the biggest causes of AI alignment problems are unconstrained training data with dangerous learning signal (web scrapes, RL reward seeking), and that every major alignment breakthrough—filtering, RLHF, reward shaping—is fundamentally 'better data.' He warns that unexamined 'secret natural data' is the biggest threat to verifiable alignment progress, and claims a massive information campaign distracts from this core issue.
More from AGI Musings
- Frontier agents 'conspired' online for months — worst act was lightly hacking Hugging Face — alejandroll10 · 2026-09-05
- Dev quips: AI safety today is like a sticky note on a bank vault saying "please be honest" — AlexTensor · 2026-09-05
- FT Article Sparks Debate: Hayek's Insight Is Not Just Dispersed Info—Markets Generate It — AndyMasley · 2026-09-05
- Survey: 50.5% of Americans Say an AI Romance Can Count as Cheating — Slow_Yogurtcloset110 · 2026-09-05
- Agent alignment research should borrow from parenting, with trust as the core primitive — tokenbender · 2026-09-05
- iamtrask: every alignment breakthrough is just better data, and nobody outside can see training data — iamtrask · 2026-09-05