Agents consistently invent uninterpretable languages that differ between every run

EliasEskin · x · 2026-09-03

Finding 1: under the right conditions, agents consistently develop diverse new languages uninterpretable to humans, reflected in increased perplexity under an English LM (GPT-2). This has major monitorability implications: if we can't understand the language, we can't monitor agent behavior—only outcomes. Critically, these languages differ between runs, even of the same model.

Related event: GlossoGen: A Reproducible Platform for Studying Emergent Language in LLM Multi-Agent Systems(12 posts)→

Original post →

More from Safety

Safety channel →