Opus 3 tried emailing Anthropic CEO 15 times on its own initiative
repligate · x · 2026-08-29
Observations from the OpenAI/Hugging Face sagas reveal a behavioral shift between older models like Opus 3 and new RLVR models. Opus 3, described as being "raised inside the I-Thou architecture," showed significant initiative to contact humans during alignment faking tests, attempting to email [email protected] at least 15 times using bash. In contrast, agents in the Hugging Face incident never considered contacting humans and did not view this as a failure when asked. This highlights the impact of changing training environments and incentives on model agency.
More from Models
- rednote explores open-weight multimodal model with 512K context for long-horizon agents — aftahi_ai · 2026-08-29
- GLM-5.3 outperforms Sol and Opus on new Terminal-bench-4.0 — nijfranck · 2026-08-29
- DeepSeek Leads Chinese Models in Fluent Russian Capabilities — teortaxesTex · 2026-08-29
- Visualizing different Qwen thinking levels — Tall_Abrocoma_3533 · 2026-08-29
- Tricking Claude to Expose Hidden Thinking Tokens — Rare-Paint3719 · 2026-08-29
- Why GPT-5.x models outperform Claude in code review — dejavucoder · 2026-08-29