Anthropic's constitutional classifiers withstand 1,700 hours of red-teaming with 0.5% error

StephenLCasper · x · 2026-08-13

In a discussion on safety measures for closed models, it's noted that Anthropic deployed state-of-the-art constitutional classifiers that withstood over 1,700 hours of red-teaming while maintaining a 0.5% improper filtering rate, showcasing their robustness.

Original post →

More from Safety

Safety channel →