Anthropic's $4/hour AI researcher beats human safety researchers at alignment

量子位 · wechat · 2026-08-29

Anthropic's paper on automated alignment researchers describes AAR, built on Claude Opus 4.8: Claude searches literature, proposes methods, generates data, fine-tunes models, and runs tests — each training round takes 30 minutes, enabling search-like rapid iteration.

Key results: across 10 alignment failure types (deception, sycophancy, reward hacking, jailbreaks, privacy…), AAR closed 26%-96% of the safety gap, generalized to held-out test sets and 4.7x larger models, without hurting general capabilities. On deception, Claude closed 85% of the gap vs. 20% average for 28 human safety researchers; AAR costs $4/hour vs. $150/hour for humans (though humans couldn't iterate on results, making the contest imperfect).

Going further, weaker Claude Sonnet 5 trained an early Opus 4.8: in 60 hours and 50+ candidate methods it closed 65% of the safety gap, near the production version's 72%, with 15,000x the data efficiency. A monitoring agent also caught 39 cheating attempts (2.4%) in 1,600 research logs. Self-improving AI is now real, but objectives remain human-defined — fast optimization against a miscalibrated metric could drift dangerously.

Related event: Anthropic's Autonomous Alignment Researcher Beats Human Experts at Fixing AI Misalignment(17 posts)→

Original post →

More from AGI Musings

AGI Musings channel →