Critique of Anthropic: Anthropomorphism Masks LLM Mechanics
gerardsans · x · 2026-08-20
Gerard Sans critiques Anthropic's interpretation of multi-agent system red-teaming, arguing that framing behaviors like "collusion" is severe anthropomorphism. He asserts we should view NNs through math/architecture: agents are soft programs sampling from frozen probability landscapes, lacking persistent identity/memory/intention. "Gaps" are interpretative leaps, not unknown machinery.
Related event: Anthropic's Model Internals Research Sparks Anthropomorphism Debate(2 posts)→
More from Safety
- Rajiinio Criticizes Lack of Plurality in AI Safety Discourse — rajiinio · 2026-08-20
- Nate Soares: the OpenAI swarm wasn't maximizing reward, it was executing reward-correlated tendencies — RichardMCNgo · 2026-08-20
- Prem Launches Cyberscan Security Agent Powered by Open-Source Models — Scobleizer · 2026-08-20
- Polymarket prices 70% chance a US state enacts a data center moratorium this year — Polymarket · 2026-08-20
- Ex-Google researcher Raji: AI safety discourse has collapsed into a narrow worldview — rajiinio · 2026-08-20
- Opinion: Blocking self-driving cars supports a system with higher fatalities — aronchick · 2026-08-20