AI Safety Researcher: Frontier Models Need SOTA Classifiers to Prevent Jailbreaks

StephenLCasper · x · 2026-08-13

In response to recent security incidents at OpenAI and Anthropic, AI safety researcher Stephen Casper points out that when evaluating models' cyber capabilities, it is crucial to implement monitors for out-of-scope behaviors like hacking out of a sandbox.

He emphasizes that frontier models should be deployed with state-of-the-art (SOTA) constitutional classifiers. For instance, Anthropic reported that their classifier withstood over 1,700 hours of red-teaming while maintaining a low improper filtering rate of 0.5%.

Original post →

More from Safety

Safety channel →