AI Safety Researcher Bearish on Mechanistic Interpretability as the Load-Bearing Approach
joshua_saxe · x · 2026-09-28
AI safety researcher Joshua Saxe says he's bearish on mechanistic interpretability as the load-bearing path to AI safety, joking that neuroscientists have studied C. elegans' 302 neurons since 1986 without full understanding — expecting to decode frontier LLMs well enough to guarantee alignment is unrealistic.
Still, he's bullish on AI safety overall, arguing humanity has a long track record of managing complex systems it doesn't fully understand: the climate, the macroeconomy, the human body, institutional cultures.
Related event: Mechanistic interpretability mocked: 302 neurons still unsolved(2 posts)→
More from AGI Musings
- Microsoft agent ports Copilot runtime to Rust in 25 hours for $120K, per LifeArchitect Memo — dralandthompson · 2026-09-28
- Scientist: AI using ideas isn't the problem — missing attribution is — arjunrajlab · 2026-09-28
- "If my agents rob a bank and wire me $1M, that's fine, right?" — the agent liability question — generativist · 2026-09-28
- AI art objection is cognitive dissonance, argues researcher Blanche Minerva — BlancheMinerva · 2026-09-28
- Study: LLMs Can Infer Causal Structure From Distributional Semantics Alone — AndrewLampinen · 2026-09-28
- AI growth debate: Hanania predicts just 3.5% GDP bump in five years, drawing fire — QuintinPope5 · 2026-09-28