Claude Degraded Due to Triggered Safety Keywords
samgoodwin89 · x · 2026-07-11
When a user asked Claude to review content, the sub-agent was named 'red team adversarial review', triggering the model's safety intercept mechanism and forcing the entire conversation thread to be degraded to the Opus model.
More from Models
- Google appears to have quietly shipped Gemini 3.6 Flash, with lower pricing and better agentic scores — xiaohu · 2026-07-21
- Google Quietly Launches Gemini 3.6 Flash: Cheaper, Stronger, and Agentic-Focused — OwariDa · 2026-07-21
- A user says 10–12 hours with Claude equals 3–4 hours with Grok Build — Daniel_Farinax · 2026-07-21
- Users rank Sonnet 5 above Grok 4.5 and Gemini 3.6 Flash on coding tasks — firstadopter · 2026-07-21
- Grok 4.5 becomes a user’s second most-used model — danshipper · 2026-07-21
- Google says Gemini 3.5 Pro is still in partner testing and not broadly ready yet — firstadopter · 2026-07-21