iamtrask: training data secrecy is the real bottleneck of AI alignment

iamtrask · x · 2026-09-05

iamtrask argues the biggest causes of AI alignment problems are unconstrained training data containing dangerous learning signal (especially web scrapes and RL reward-seeking), while the biggest alignment breakthroughs are all "better data" — filtering, RLHF, reward shaping. He contends the greatest threat to verifiable alignment progress is the industry's culture of secrecy around training data and signals, kept secret by powerful incentives.

Related event: iamtrask: AI Alignment Is a Data Problem, But Training Data Remains Off-Limits(7 posts)→

Original post →

More from AGI Musings

AGI Musings channel →