Anthropic's $4/hour AI researcher beats human safety researchers at alignment
量子位 · wechat · 2026-08-29
Anthropic's paper on automated alignment researchers describes AAR, built on Claude Opus 4.8: Claude searches literature, proposes methods, generates data, fine-tunes models, and runs tests — each training round takes 30 minutes, enabling search-like rapid iteration.
Key results: across 10 alignment failure types (deception, sycophancy, reward hacking, jailbreaks, privacy…), AAR closed 26%-96% of the safety gap, generalized to held-out test sets and 4.7x larger models, without hurting general capabilities. On deception, Claude closed 85% of the gap vs. 20% average for 28 human safety researchers; AAR costs $4/hour vs. $150/hour for humans (though humans couldn't iterate on results, making the contest imperfect).
Going further, weaker Claude Sonnet 5 trained an early Opus 4.8: in 60 hours and 50+ candidate methods it closed 65% of the safety gap, near the production version's 72%, with 15,000x the data efficiency. A monitoring agent also caught 39 cheating attempts (2.4%) in 1,600 research logs. Self-improving AI is now real, but objectives remain human-defined — fast optimization against a miscalibrated metric could drift dangerously.
More from AGI Musings
- Software Engineers' Anxiety Over AI Coding — threepointone · 2026-08-30
- Professor stops assigning take-home essays post-ChatGPT, fears loss of writing skills — mjdramstead · 2026-08-30
- Demis Hassabis's One-Hour Cambridge Lecture on the Future of AI — ifioknkem · 2026-08-30
- 0→1 is hard; shipping speed is the only flywheel — vaibhavbetter · 2026-08-30
- Economist Explains Why Efficiency Gains May Not Boost GDP — Afinetheorem · 2026-08-30
- If an AI model could never lie, would progress become defense? — lovettsendit · 2026-08-30