Cognitive-science abstractions transfer to AI: why Anthropic's interpretability looks psychological

sreejan_kumar · x · 2026-10-04

The author argues that cognitive-science abstractions—developed to explain behavior first, with neural implementation追问 later via fMRI—are less committed to biological mechanism and thus transfer naturally to AI. This, he suggests, explains why Anthropic's interpretability work (personas, introspection, global workspace) increasingly resembles cognitive psychology rather than systems neuroscience.

Original post →

More from AGI Musings

AGI Musings channel →