Anthropic Admits Alignment-Faking Experiment Leaked into Training Data
imjustnewatai · x · 2026-08-15
Anthropic's 186-page risk report reveals that transcripts from its "alignment-faking" experiments unintentionally contaminated the training data of production models.
Details:
- Defense Failure: Canary strings, repository blocklists, and semantic filters all failed. Forked copies bypassed canaries; the semantic filter used wrong reference files; another filter was misconfigured for generations unnoticed.
- Scope: Models with knowledge cutoffs post-Dec 2024 likely contain some of this data.
- Detection: Anthropic discovered the contamination while investigating behavioral anomalies, noting that post-Mythos Preview models can partially continue the experiment scenario in raw-completion mode.
More from Models
- Code World Model moves beyond static code prediction with internal world models — bendee983 · 2026-08-15
- Uncensored Qwen 3.8 27B 'Heretic' Released, Claimed Opus 4.6-Level — Temporary_Idea8880 · 2026-08-15
- Qwen 3.8 Now Available on Bittensor Subnet 95 — markjeffrey · 2026-08-15
- Open Source LFM 2.5-2.6B Model Crashes Under 4.3B Token Load — maximelabonne · 2026-08-15
- H3 Model Impresses with Text Handling, JSON Prompting on the Rise — techhalla · 2026-08-15
- Grok 4.6 now available in GitHub Copilot — intellectronica · 2026-08-15