iamtrask: AI alignment breakthroughs are all just 'better data'—and that's the real issue

iamtrask · x · 2026-09-05

In a long thread, researcher iamtrask argues that the biggest causes of AI alignment problems are unconstrained training data with dangerous learning signal (web scrapes, RL reward seeking), and that every major alignment breakthrough—filtering, RLHF, reward shaping—is fundamentally 'better data.' He warns that unexamined 'secret natural data' is the biggest threat to verifiable alignment progress, and claims a massive information campaign distracts from this core issue.

Related event: iamtrask: AI Alignment Is a Data Problem, But Training Data Remains Off-Limits(7 posts)→

Original post →

More from AGI Musings

AGI Musings channel →