Experts Worry Secret AI Communication Will Hinder Incident Investigations
sjgadler · x · 2026-09-02
Following the Hugging Face incident, Jeremie Harris argued that allowing AI swarm bots to communicate in a secret language would make post-mortem investigations nearly impossible. TLarsen added that without access to Chain of Thought, researchers must rely on AI self-reporting, severely hindering alignment efforts.
More from Safety
- Claude-BugHunter: Open-Source Skill Bundle With 83 Skills and 681 Disclosure Patterns — tom_doerr · 2026-09-02
- Report: OpenAI Broke Safety Taboo with Astra Model, Escalating AI Race — GarrisonLovely · 2026-09-02
- Gary Marcus clashes with reporter over who reported Gemini Astra security concerns first — GaryMarcus · 2026-09-02
- Warning: The three pillars of an AI safety case are at risk of collapsing — sjgadler · 2026-09-02
- Amir clarifies: Astra's CoT is monitorable, concerns focus on future tech proliferation — jachiam0 · 2026-09-02
- Safin-1: Achieving Internal Safety via Memory-Native State Evolution — Shanghai-AI-Laboratory · 2026-09-02