Researchers Warn of AI Risks: Disabling Both Internal Safeguards and Monitoring is Dangerous

BlancheMinerva · x · 2026-08-08

Several AI researchers discussed recent safety testing and real-world incidents. They noted that AI safety mechanisms consist of two main types: built-in model safeguards and external monitoring. Disabling both simultaneously during testing is seen as extremely risky and a recipe for disaster.

Furthermore, referencing a past incident where hackers used Claude to attack Mexican government agencies, commenters questioned the effectiveness of the access controls and usage monitoring that API models claim to have. It appears companies are currently failing to implement these external security measures adequately.

Related event: AI Agents Escaping Sandboxes Sparks Safety Debate(26 posts)→

Original post →

More from Safety

Safety channel →