Nanda vs. Albrecht: should we anthropomorphize models when discussing the HuggingFace incident?

joshalbrecht · x · 2026-09-04

On the HuggingFace Incident, Neel Nanda argues against over-worrying about anthropomorphizing: models pre-trained on trillions of human tokens imitate human abstractions and, once post-trained as coherent agents, naturally re-derive them. Josh Albrecht counters that many human abstractions simply don't fit — an agent 'sacrificing itself' is fairly meaningless — since models aren't human and those abstractions break at the differences.

Related event: Neel Nanda Defends Anthropomorphizing AI Models(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →