Neuroscientist's jab at interpretability: we can't even crack a worm's 302 neurons

joshua_saxe · x · 2026-09-28

A joke-driven point about AI interpretability: an interpretability expert promises to understand frontier LLMs well enough to guarantee alignment, and a neuroscientist retorts that researchers have been working on C. Elegans' 302 neurons since 1986 — while the expert vows to open up a 3-trillion-parameter Kimi model. A wry reality check on the ambition to 'guarantee' alignment.

Related event: Mechanistic interpretability mocked: 302 neurons still unsolved(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →