iamtrask: AI Alignment Is a Data Problem, But Training Data Remains Off-Limits
iamtrask (Andrew Trask), a researcher with roots at OpenAI, posted a series of tweets on September 5 laying out his core argument: AI alignment is fundamentally a data problem, and the biggest obstacle facing alignment research is that researchers cannot access AI companies' most closely guarded asset—training data.
Confirmed
- iamtrask argues that everything an AI knows, believes, and does comes from training and reinforcement learning data, so the biggest cause of and the biggest solution to alignment both point to the same place: data.
- He identifies the main "pathology" of misalignment as dangerous learning signals embedded in unconstrained training data—especially web-scraped corpora and reward-seeking behavior during RL.
- He notes that nearly every breakthrough in alignment has come from "better data"—techniques like data filtering, RLHF, and reward shaping.
- He stresses that AI companies' most confidential asset is neither model weights nor user logs, but training (including RL) data; denying alignment researchers unrestricted access to these data records is like being asked to find an antidote without being allowed to study the poison—progress is possible, but far harder.
- He admits the topic is "deeply unsexy," but believes the argument is straightforward.
Why it matters
- This view shifts the transparency debate in alignment research toward training data—typically kept under strict secrecy—rather than the more commonly discussed model weights or user logs, raising a new demand for openness from external oversight and safety research.
- If his argument holds, "better data" becomes the main thrust of alignment work, and data governance and access rights could become central issues in AI safety policy.
2026-09-05 ~ 2026-09-05 · 7 related posts
Primary sources
- Trask: alignment research without training data access is like finding antidotes blind — iamtrask ·
- iamtrask: training data secrecy is the real bottleneck of AI alignment — iamtrask ·
- iamtrask: alignment may only be solved by the few with access to AI's most secret asset — training data — iamtrask ·
- iamtrask: every alignment breakthrough is just better data, and nobody outside can see training data — iamtrask · 2026-09-05
- iamtrask: AI alignment breakthroughs are all just 'better data'—and that's the real issue — iamtrask · 2026-09-05
- [source] iamtrask: training data secrecy is the real bottleneck of AI alignment — iamtrask · 2026-09-05
- iamtrask: AI alignment is a data problem — every breakthrough is just better data — iamtrask · 2026-09-05
- iamtrask: alignment is a data problem — without training data access you're in the dark — iamtrask · 2026-09-05
- [source] Trask: alignment research without training data access is like finding antidotes blind — iamtrask · 2026-09-05
- [source] iamtrask: alignment may only be solved by the few with access to AI's most secret asset — training data — iamtrask · 2026-09-05