Anthropic: Models can automatically improve safety benchmarks without degrading capabilities

AnthropicAI · x · 2026-08-29

Anthropic had Claude "hill-climb" safety benchmarks for common misalignments like deception or sycophancy, with the constraint of preserving general capabilities. The best methods generalized to held-out benchmarks and reliably improved safety scores.

Related event: Anthropic Shows Claude Can Automatically Align Other Models(3 posts)→

Original post →

More from Safety

Safety channel →