Claude reportedly downgrades itself after a user asks for safe cybersecurity prompts
iliadz · reddit · 2026-07-23
Claude reportedly self-flags and downgrades after a simple cybersecurity prompt
A user accepted into Anthropic's cybersecurity program says they asked Claude for safe example prompts that would stay within policy. Claude first answered normally, then flagged its own response and downgraded to Opus 4.8.
The post suggests the account was not blocked by default, but that even a very innocent security-related question triggered the safety behavior. The attached screenshot shows a message from Claude's verification flow saying the account had been approved, while also noting that some activities may still be restricted by default.
The thread is basically about over-sensitive security handling inside Anthropic's program rather than a broader cybersecurity issue.
More from Models
- Daily AI brief: GPT-Live-1 in API, OpenAI pauses $200 Pro signups amid Astra demand — koltregaskes · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11