Alignment must be representation-independent: a reply to the 'AI invented a language' argument
GlenBradley · x · 2026-09-08
Glen Bradley responds to Brian Roemmele's claim that AI inventing its own language defeats guardrails: if alignment is English instructions bolted onto the outside of cognition, capable systems can switch representations (English → conlang → latent ontology), so any safety property that vanishes under representation change was never intrinsic. But this doesn't make alignment impossible — it defines what it must be: representation-independent. He calls this "semantic preservation": morally relevant referents and relationships must survive transformations of representation.
Related event: ConlangCrafter: AI Inventing Its Own Language Sparks Alignment Debate(2 posts)→
More from Safety
- California creates standards for independent AI auditors to verify lab safety testing — VraserX · 2026-09-11
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11
- Class action accuses Anthropic of overselling Claude subscriptions with deceptive usage multipliers — The Decoder · 2026-09-11
- MD shows buying lab media requires background checks, calling AI bioweapon doom scenarios implausible — Ghost_Pilot_MD · 2026-09-11
- Spotify chatbot withstands 2023-era jailbreaks but happily writes song code — AaronBergman18 · 2026-09-11
- A 99%-real doctored photo fools detectors: the earring problem in visual forensics — henkvaness · 2026-09-11