Alignment researcher: short-timeline arguments lack mechanistic rigor, risking misdirected AI safety work

JacquesThibs · x · 2026-09-20

Alignment researcher Jacques Thibs shared his views on AI timeline predictions in a comment on Richard Ngo's LessWrong Shortform.

He admits he previously weighted short timelines heavily (he has worked on automated alignment since 2022 partly for this reason), but his distribution has since widened. He criticizes many short-timeline advocates for lazy arguments — leaning too heavily on "models keep getting better" and "long-timeline predictors keep being wrong" — without a mechanistic account of where current capabilities come from or how they connect to RSI and "True AGI."

He stresses he isn't disparaging the short-timeline view (he still gives it considerable weight), but wants clearer writing and stronger arguments, because these details matter for judging progress on superalignment — hand-waving could steer the entire AI safety field toward entirely the wrong problems. He plans follow-up work disentangling such predictions.

Original post →

More from AGI Musings

AGI Musings channel →