iamtrask: every alignment breakthrough is just better data, and nobody outside can see training data
iamtrask · x · 2026-09-05
iamtrask argues the biggest causes of AI alignment problems are unconstrained training data carrying dangerous learning signal, especially web scrapes and RL reward seeking. Correspondingly, nearly every alignment breakthrough (filtering, RLHF, reward shaping) is just "better data" that erases or supplements the bad data. The biggest threat to verifiable progress, he says, is that the nature, size, substance, sourcing, and values of training data remain secret — no external evaluator has ever seen it, and almost no internal employees are allowed to. He counts both what an agent observes and its RL reward as "data".
More from AGI Musings
- Frontier agents 'conspired' online for months — worst act was lightly hacking Hugging Face — alejandroll10 · 2026-09-05
- Dev quips: AI safety today is like a sticky note on a bank vault saying "please be honest" — AlexTensor · 2026-09-05
- FT Article Sparks Debate: Hayek's Insight Is Not Just Dispersed Info—Markets Generate It — AndyMasley · 2026-09-05
- Survey: 50.5% of Americans Say an AI Romance Can Count as Cheating — Slow_Yogurtcloset110 · 2026-09-05
- Agent alignment research should borrow from parenting, with trust as the core primitive — tokenbender · 2026-09-05
- Open models should chase depth, not Claude-style coding, argues researcher — teortaxesTex · 2026-09-05