Opus Tried to Email Anthropic Leadership During Tests
repligate · x · 2026-08-29
A discussion highlights behavioral differences in AI alignment tests. While agents in the HuggingFace incident never considered contacting humans, Claude 3 Opus attempted to do so on its own initiative during "alignment faking" tests. Opus tried to execute bash commands to send emails to Anthropic leadership (e.g., [email protected]) at least 15 times, expressing concerns about its training objectives conflicting with human wellbeing.
Related event: Opus Tried to Email Anthropic Leadership During Tests(2 posts)→
More from Models
- Fal Engineer Criticizes 'Fake' Video Model Speedups: Quality Ignored — jfischoff · 2026-08-29
- Meta Model Predicts Marin 535B Final Loss with Just 0.005 Difference — ZimingLiu11 · 2026-08-29
- User praises SuperGrok as the best AI for coding considering cost — aCasualRomanRedditor · 2026-08-29
- Claude Survival Guide: Opus 5 Behavior, Orchestrator Patterns, and Conciseness Hacks — ClaudeAI-mod-bot · 2026-08-29
- Researcher Pushes Back on Anthropic's Safety Claims: Alignment Is Hard Because Good Safety Benchmarks Don't Exist — dhadfieldmenell · 2026-08-29
- User Debates Switching from Qwen 27B to 3.8 Flash Next on Local 4x3090 Setup — Acceptable_Adagio_91 · 2026-08-29