Claude safety filters keep escalating a harmless shell command
P_nde · reddit · 2026-07-25
A screenshot shows Claude’s safety system repeatedly flagging a harmless shell command, escalating from Fable 5 to Opus 5, then Sonnet, and finally Haiku as each layer complains about the previous one.
The joke is that the guardrails are so sensitive that the model keeps being downgraded until the smallest version is left to answer, turning a safety feature into a comic chain reaction.
More from Fun
- Tesla FSD blamed for crossing floating bridge at 75 MPH — a Chevrolet was actually the culprit — mariolefebvre · 2026-09-11
- X drama: Anthropic researchers accused of spying on academic customers and racing them to results — basedjensen · 2026-09-11
- Llama 405B's Dark Inventions Creep Out Opus in an AI Word Game — liminal_bardo · 2026-09-11
- fable 5.1 recreates The Starry Night with 256,157 JavaScript brush strokes — cedric_chee · 2026-09-11
- "Anyone still coding the old way?" The joke capturing post-AI programming culture — lxfater · 2026-09-11
- iLands agents email philosopher asking $20 for piecework, sparking unease about AI consciousness — tobyordoxford · 2026-09-11