User Reports Claude Safety Classifier Triggering Loop Errors

East_Trust_9588 · reddit · 2026-07-05

A user reported experiencing false triggers from Claude's safety classifier: simply using the word "ready" in a non-suicide-related context caused the model to repeatedly interrupt the response and pivot back to discussing suicide.

The user pointed out that even if the classifier misfires, the system prompt's mandate to 'address it directly' forces the model to loop back to the topic. This could be harmful to those actually seeking help, prompting a call for Anthropic to improve the mechanism.

Original post →

More from Models

Models channel →