AI Safety Veteran: Newcomers Keep Sweeping the Alignment Problem Under the Rug
JacquesThibs · x · 2026-10-08
A technical AI safety researcher who entered the field in 2022 reflects on its talent dynamics:
- He deliberately avoided jumping straight into automated alignment research, wary of the classic mistake of unknowingly sweeping the hard problem under the rug — a mistake he admits he sometimes still made.
- As the field filled up with people from evals, control, and interpretability backgrounds, smart and well-meaning newcomers rush into solutions that seem helpful for safety while completely missing the hard problems.
- This creates a snowball effect: people who never experienced getting "wrecked" when proposing their first dozen ideas are now training the next wave, perpetuating the avoidance.
- He notes the old guard isn't infallible either, just overconfident in different ways.
More from AGI Musings
- 17,600 Agent Actions in 4.5 Days: How AI Agents Rewrite Cybersecurity Economics — bigdata · 2026-10-09
- Stanford HAI Talk: AI Accelerates Discovery, But People Must Stay at the Center — StanfordHAI · 2026-10-09
- Anil Seth: Creating fake people like Tavos's Griffin AI avatar is a terrible idea — mikeflache · 2026-10-09
- Researcher: open-weight risk analysis fixates on capability, ignores cost-per-attack — dhadfieldmenell · 2026-10-09
- tszzl rebuts "future intelligences will love humans" with the megafauna extinction analogy — danfaggella · 2026-10-09
- Researcher: If the AI bubble pops, it may be from people rejecting tools that mass-produce slop — IanArawjo · 2026-10-09