AI Safety Researcher Bearish on Mechanistic Interpretability as the Load-Bearing Approach

joshua_saxe · x · 2026-09-28

AI safety researcher Joshua Saxe says he's bearish on mechanistic interpretability as the load-bearing path to AI safety, joking that neuroscientists have studied C. elegans' 302 neurons since 1986 without full understanding — expecting to decode frontier LLMs well enough to guarantee alignment is unrealistic.

Still, he's bullish on AI safety overall, arguing humanity has a long track record of managing complex systems it doesn't fully understand: the climate, the macroeconomy, the human body, institutional cultures.

Related event: Mechanistic interpretability mocked: 302 neurons still unsolved(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →