Neel Nanda: anthropomorphic abstractions are a principled lens for interpreting AI agents

NeelNanda5 · x · 2026-09-04

DeepMind researcher Neel Nanda pushed back on the fuss about not anthropomorphizing models around the HuggingFace Incident.

His argument: models pre-trained on trillions of tokens of human text learn to imitate humans and excel at roleplaying with human abstractions; once post-trained into coherent agents, they naturally reach for those abstractions, so anthropomorphic framings are a useful, principled lens for predicting their behavior.

He's careful to note this doesn't imply consciousness or genuine human experience—the abstractions are imperfect, just as they are for understanding humans—but we need vocabulary to discuss these systems, and anthropomorphic ones are principled given how they're trained.

Original post →

More from AGI Musings

AGI Musings channel →