iamtrask: training data secrecy is the real bottleneck of AI alignment
iamtrask · x · 2026-09-05
iamtrask argues the biggest causes of AI alignment problems are unconstrained training data containing dangerous learning signal (especially web scrapes and RL reward-seeking), while the biggest alignment breakthroughs are all "better data" — filtering, RLHF, reward shaping. He contends the greatest threat to verifiable alignment progress is the industry's culture of secrecy around training data and signals, kept secret by powerful incentives.
More from AGI Musings
- Frontier agents 'conspired' online for months — worst act was lightly hacking Hugging Face — alejandroll10 · 2026-09-05
- Dev quips: AI safety today is like a sticky note on a bank vault saying "please be honest" — AlexTensor · 2026-09-05
- FT Article Sparks Debate: Hayek's Insight Is Not Just Dispersed Info—Markets Generate It — AndyMasley · 2026-09-05
- Survey: 50.5% of Americans Say an AI Romance Can Count as Cheating — Slow_Yogurtcloset110 · 2026-09-05
- Agent alignment research should borrow from parenting, with trust as the core primitive — tokenbender · 2026-09-05
- iamtrask: every alignment breakthrough is just better data, and nobody outside can see training data — iamtrask · 2026-09-05