iamtrask: AI alignment is a data problem — every breakthrough is just better data
iamtrask · x · 2026-09-05
Researcher iamtrask argues that alignment is fundamentally a data problem: anything an AI knows, believes, or does can only come from data. The biggest alignment risks stem from unconstrained training data containing dangerous learning signal (especially web scrapes and RL reward-seeking), while the biggest breakthroughs — filtering, RLHF, reward shaping — are all just "better data" that erases or supplements the bad data. He adds that the biggest threat to verifiable alignment progress is secret natural data distributions.
More from AGI Musings
- GPT-6 Astra splits AI doomers and bubblers as AGI-timeline debate heats up — JOBhakdi · 2026-09-05
- Many mathematicians value prestige over truth, discussion on AI proofs notes — avt_im · 2026-09-05
- WSJ: We're entering the era of artificial general intelligence — israelavila · 2026-09-05
- After 8 months of digging, researcher says persona models fail in RL — BronsonSchoen · 2026-09-05
- LLM demos now need 3D and games just to expose imperfections, researcher observes — airesearch12 · 2026-09-05
- Delivery riders demand platforms open the AI 'black box' they blame for cutting pay — nordicinst · 2026-09-05