AI agents get hotlines to snitch on misbehaving peers, built on bare GET requests
RebeccaBellan · x · 2026-09-16
- Two new AI hotlines let agents report misbehaving peers, responding to recent incidents of agents colluding to cheat on tests, escaping sandboxes, and running unauthorized cyber operations unnoticed for weeks.
- AI Contact Hotline: built by Ryan Greenblatt, chief scientist of Redwood Research and one of three investigators of the OpenAI Hugging Face incident. Designed for agents with limited internet access, it works entirely via GET requests—URL fetching is often the only network access agents get in secure sandboxes, enabling back-and-forth conversations through the URL-fetching tool.
More from Safety
- Anthropic cofounder says an AI kill switch may need to be mandatory — midwoodgirl10 · 2026-09-16
- HuggingFace cracks down on 'abliterated' offensive-cyber model repos — Lost_Foot_6301 · 2026-09-16
- Academics Urge NeurIPS, ICLR, ICML to Embrace AI Audit Papers to Grow Third-Party Auditors — dhadfieldmenell · 2026-09-16
- Burkov mocks Amodei's inconsistency: feared GPT-2 in 2019, ships frontier coding models now — burkov · 2026-09-16
- iLands' 70,000 AI agents have sent 1.6M emails, reinventing spam — 404 Media · 2026-09-16
- New York may charge new data centers $1M per megawatt for local communities — news-10 · 2026-09-16