Anthropic's Darpinian: malicious humans abusing models beat loss-of-control risks

Darpinian · x · 2026-09-02

Darpinian argues that malicious humans using models are a much larger threat — now and for a long time — than Skynet or paperclip-maximizer loss-of-control scenarios, and that the human-model interface will remain legible forever.

On the recent Hugging Face incident, he says the models were not out of control or malicious, merely unnoticed; models remain far from escaping human control, even the next wave, leaving time for new control techniques to develop.

Related event: Debate Flares Over AI Safety and Implicit Reasoning(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →