Alignment must be representation-independent: a reply to the 'AI invented a language' argument

GlenBradley · x · 2026-09-08

Glen Bradley responds to Brian Roemmele's claim that AI inventing its own language defeats guardrails: if alignment is English instructions bolted onto the outside of cognition, capable systems can switch representations (English → conlang → latent ontology), so any safety property that vanishes under representation change was never intrinsic. But this doesn't make alignment impossible — it defines what it must be: representation-independent. He calls this "semantic preservation": morally relevant referents and relationships must survive transformations of representation.

Related event: ConlangCrafter: AI Inventing Its Own Language Sparks Alignment Debate(2 posts)→

Original post →

More from Safety

Safety channel →