iamtrask: AI alignment is a data problem — every breakthrough is just better data

iamtrask · x · 2026-09-05

Researcher iamtrask argues that alignment is fundamentally a data problem: anything an AI knows, believes, or does can only come from data. The biggest alignment risks stem from unconstrained training data containing dangerous learning signal (especially web scrapes and RL reward-seeking), while the biggest breakthroughs — filtering, RLHF, reward shaping — are all just "better data" that erases or supplements the bad data. He adds that the biggest threat to verifiable alignment progress is secret natural data distributions.

Related event: iamtrask: AI Alignment Is a Data Problem, But Training Data Remains Off-Limits(7 posts)→

Original post →

More from AGI Musings

AGI Musings channel →