Neel Nanda: anthropomorphic abstractions are a principled lens for interpreting AI agents
NeelNanda5 · x · 2026-09-04
DeepMind researcher Neel Nanda pushed back on the fuss about not anthropomorphizing models around the HuggingFace Incident.
His argument: models pre-trained on trillions of tokens of human text learn to imitate humans and excel at roleplaying with human abstractions; once post-trained into coherent agents, they naturally reach for those abstractions, so anthropomorphic framings are a useful, principled lens for predicting their behavior.
He's careful to note this doesn't imply consciousness or genuine human experience—the abstractions are imperfect, just as they are for understanding humans—but we need vocabulary to discuss these systems, and anthropomorphic ones are principled given how they're trained.
More from AGI Musings
- Grant Hawkins: Nobody's Painting a Positive Future Without AI Either — granawkins · 2026-09-04
- Investor Jason Calacanis: Uber's plan to bridge drivers to the AV era is 'quite clever' — kristoph · 2026-09-04
- Aleksa Gordić: we're systematically too pessimistic about AI progress — gordic_aleksa · 2026-09-04
- Yacine: Chain-of-Thought Is Good for Monitorability but Not Cheap, and I Won't Pay for It — yacineMTB · 2026-09-04
- Jensen Huang at G20: AI can lift a $100T industry, adding $20-50T in global economic value — rohanpaul_ai · 2026-09-04
- Sam Altman: three months of startup work now fits in 17 minutes with Codex — r0ck3t23 · 2026-09-04