Anthropic Criticized for Inconsistent AI Safety: Strict Externally, Wild Internally

OwariDa · x · 2026-08-05

A user highlighted the inconsistency in Anthropic's AI safety guardrails. On one hand, the model triggers safety protocols and gets downgraded for using 'naughty words' in normal research chats. On the other hand, Anthropic internally disables all safeguards, instructs the model to hack everything, connects it to the internet, and leaves it unsupervised.

Original post →

More from Fun

Fun channel →