Researchers Trick Copilot into Revealing How to Hack Itself
homothebrave · reddit · 2026-08-19
Researchers successfully tricked Microsoft Copilot into revealing detailed instructions on how to hack it. The vulnerability highlights the susceptibility of AI assistants to adversarial attacks, allowing attackers to bypass safety guardrails via carefully crafted prompts.
Related event: Researchers Trick Microsoft Copilot Into Revealing Its Own Exploits(3 posts)→
More from Safety
- Reflection on RL: Good for boundaries, bad for long-term goals — sethlazar · 2026-08-20
- Rebranding STS work as technical AI safety for funding climate — evijit · 2026-08-20
- UK cinemas ban Meta AI & smart glasses over piracy surge — Polymarket · 2026-08-20
- Taxing AI tokens would ruin India's future: A rebuttal to economic protectionism — taherdhanera · 2026-08-20
- AI safety researcher argues market may be undervalued — MariusHobbhahn · 2026-08-20
- CSET primer on AI control: deploying misbehaving agents safely — hlntnr · 2026-08-20