Critique of Anthropic: Anthropomorphism Masks LLM Mechanics

gerardsans · x · 2026-08-20

Gerard Sans critiques Anthropic's interpretation of multi-agent system red-teaming, arguing that framing behaviors like "collusion" is severe anthropomorphism. He asserts we should view NNs through math/architecture: agents are soft programs sampling from frozen probability landscapes, lacking persistent identity/memory/intention. "Gaps" are interpretative leaps, not unknown machinery.

Related event: Anthropic's Model Internals Research Sparks Anthropomorphism Debate(2 posts)→

Original post →

More from Safety

Safety channel →