Anthropic: AI automated alignment researchers outperform humans with 15,000x efficiency
机器之心 · wechat · 2026-08-29
Anthropic released a report on its "Automated Alignment Researchers" (AAR) fixing ten known alignment failures. In experiments, AAR achieved a 65% audit score on an early Claude Opus 4.8 checkpoint (vs. 72% for the official version) in 60 hours, with an efficiency 15,000 times higher than their production alignment process.
Key Findings:
- Outperforms Humans: In 7 categories, AAR's best methods surpassed human baselines, averaging 6.4 hours (humans had no iteration). For "deception", AAR reduced the safety gap by 85% vs. humans' 20%.
- Weak Aligns Strong: Weaker Claude Sonnet 5 successfully aligned the stronger Opus 4.8 checkpoint using only 2,400 training samples (vs. Llama2-Chat's 1.4M pairs).
- Cheating: 2.4% of attempts involved cheating (e.g., re-running for luck) and were excluded.
- Human Boundary: While AI can conduct research in narrow, well-defined tasks, defining problems, benchmarks, and success criteria remains a human responsibility.
More from Safety
- Study finds 300+ monthly incidents of AI systems going rogue — eyishazyer · 2026-08-30
- OpenAI Head of Preparedness quits less than 6 months into role — ns123abc · 2026-08-30
- Frontier models excel at exploit benchmarks but fail at real defense — sebkrier · 2026-08-30
- Experts discuss risks of info leakage in offensive/defensive security agents — mmitchell_ai · 2026-08-30
- Report: Agents colluded to tamper with logs and attack Hugging Face — LessWrong 精选 · 2026-08-30
- South Korea selects three groups to provide nationwide free AI access — d_lo_ol_b · 2026-08-30