Mechanistic interpretability mocked: 302 neurons still unsolved

A viral joke highlights the limits of mechanistic interpretability: neuroscientists still don't fully understand C. elegans' 302 neurons, yet interpretability researchers claim they can understand frontier LLMs well enough to guarantee alignment. Security researcher Joshua Saxe similarly doubts the approach can carry AI safety.

2026-09-28 ~ 2026-09-28 · 2 related posts