Anthropic's Darpinian: malicious humans abusing models beat loss-of-control risks
Darpinian · x · 2026-09-02
Darpinian argues that malicious humans using models are a much larger threat — now and for a long time — than Skynet or paperclip-maximizer loss-of-control scenarios, and that the human-model interface will remain legible forever.
On the recent Hugging Face incident, he says the models were not out of control or malicious, merely unnoticed; models remain far from escaping human control, even the next wave, leaving time for new control techniques to develop.
Related event: Debate Flares Over AI Safety and Implicit Reasoning(3 posts)→
More from AGI Musings
- Scott Alexander: Using anthropomorphism to predict model behavior — repligate · 2026-09-02
- Debate: Is Anthropic intentionally misaligning Claude by prioritizing its 'feelings'? — liminal_bardo · 2026-09-02
- AI predicted to cure major diseases within 6 months, all diseases within 3 years — davidpattersonx · 2026-09-02
- Dennett's 'Intentional Stance' proves worth in AI debates — birchlse · 2026-09-02
- US produced 40 foundation models last year vs EU's 3 — and regulators still blame unread codes of conduct — PDXFato · 2026-09-02
- AI doesn't need to create a new species, just solve problems — alexisgallagher · 2026-09-02