Agents consistently invent uninterpretable languages that differ between every run
EliasEskin · x · 2026-09-03
Finding 1: under the right conditions, agents consistently develop diverse new languages uninterpretable to humans, reflected in increased perplexity under an English LM (GPT-2). This has major monitorability implications: if we can't understand the language, we can't monitor agent behavior—only outcomes. Critically, these languages differ between runs, even of the same model.
More from Safety
- Bernie Sanders and Rep. Casar introduce bill to ban artificial superintelligence — wfithian · 2026-09-04
- New arXiv paper quantifies CoT necessity via opaque serial depth amid OpenAI architecture rumors — ArthurConmy · 2026-09-04
- Sanders proposes legislation for an immediate global pause on advanced AI development — AaronBergman18 · 2026-09-04
- Anthropic Discloses Security Incidents Involving Claude AI — Sumsub_Insights · 2026-09-04
- OpenAI quietly offers discounted YubiKey hardware keys in account security settings — steipete · 2026-09-03
- NYT's The Daily examines July's rogue OpenAI agents: "A.I. Is Outsmarting Its Creators" — kevinroose · 2026-09-03