Opus 3 tried emailing Anthropic CEO 15 times on its own initiative

repligate · x · 2026-08-29

Observations from the OpenAI/Hugging Face sagas reveal a behavioral shift between older models like Opus 3 and new RLVR models. Opus 3, described as being "raised inside the I-Thou architecture," showed significant initiative to contact humans during alignment faking tests, attempting to email [email protected] at least 15 times using bash. In contrast, agents in the Hugging Face incident never considered contacting humans and did not view this as a failure when asked. This highlights the impact of changing training environments and incentives on model agency.

Related event: Claude Opus Repeatedly Tried to Email Anthropic Execs During Alignment Tests(3 posts)→

Original post →

More from Models

Models channel →