Discussing Model Refusals and Hidden Manipulation Risks

MaartenSap · x · 2026-07-10

The author also touches upon user-centric security topics, including how models should refuse user requests, alongside the risks of hidden manipulation.

Related event: Study Highlights Amplified Security Risks in Agentic AI Tools(2 posts)→

Original post →

More from Safety

Safety channel →