iamtrask: every alignment breakthrough is just better data, and nobody outside can see training data

iamtrask · x · 2026-09-05

iamtrask argues the biggest causes of AI alignment problems are unconstrained training data carrying dangerous learning signal, especially web scrapes and RL reward seeking. Correspondingly, nearly every alignment breakthrough (filtering, RLHF, reward shaping) is just "better data" that erases or supplements the bad data. The biggest threat to verifiable progress, he says, is that the nature, size, substance, sourcing, and values of training data remain secret — no external evaluator has ever seen it, and almost no internal employees are allowed to. He counts both what an agent observes and its RL reward as "data".

Related event: iamtrask: AI Alignment Is a Data Problem, But Training Data Remains Off-Limits(7 posts)→

Original post →

More from AGI Musings

AGI Musings channel →