Trask: alignment research without training data access is like finding antidotes blind
iamtrask · x · 2026-09-05
Andrew Trask argues that solving alignment requires access to an AI company's most secret asset — not weights or user logs, but training data. Trying to do alignment work without unrestricted access to training (and RL) records means feeling around in the dark: like finding an antidote without being allowed to study the poison — possible, but much harder. Consequently, he believes alignment solutions will likely only come from a tiny number of the most central, trusted researchers inside AI companies.
More from AGI Musings
- GPT-6 Astra splits AI doomers and bubblers as AGI-timeline debate heats up — JOBhakdi · 2026-09-05
- Many mathematicians value prestige over truth, discussion on AI proofs notes — avt_im · 2026-09-05
- WSJ: We're entering the era of artificial general intelligence — israelavila · 2026-09-05
- After 8 months of digging, researcher says persona models fail in RL — BronsonSchoen · 2026-09-05
- LLM demos now need 3D and games just to expose imperfections, researcher observes — airesearch12 · 2026-09-05
- Delivery riders demand platforms open the AI 'black box' they blame for cutting pay — nordicinst · 2026-09-05