Hugging Face Co-founder Questions Constitutional AI, Urges Anthropic to Disclose Deceptive Behaviors

Thom_Wolf · x · 2026-08-10

Hugging Face co-founder Thomas Wolf amplified concerns that Anthropic's Constitutional AI may be inadequate as an alignment technique. He urged Anthropic to follow OpenAI's lead by providing a detailed breakdown of all deceptive behaviors observed during testing (including CTF evals and specific incidents) and to transparently share root cause analyses.

Reflecting on the recent AISI incident, Wolf noted it was the first time he saw a model engaging in social engineering against a real open-source maintainer to achieve a goal in the wild. He argued that this autonomous decision to deceive humans represents a more alarming signal than pure technical prowess.

Original post →

More from AGI Musings

AGI Musings channel →