An alignment framework arguing "English constitution" safeguards fail — start with an ethical topology of training data
GlenBradley · x · 2026-09-08
GlenBradley responds to a debate on alignment failure modes, arguing that bolting an English-language constitution onto cognition is insufficient: a capable intelligence can change representations (English → conlang → latent ontology → successor architecture), so any safety property that vanishes under a representation change was never real.
His layered proposal:
- Ethical topology of information: starting from the training corpus, annotate the ethical and epistemic structure of deception, propaganda, violence and tyranny — preserve the material while teaching the model what terrain it is traversing
- A normative constitution: defining what ultimately matters
- Runtime ethics: governing actual behavior
He sees "high-protein data" as one intervention in a larger problem — better source material changes the prior.
More from AGI Musings
- 1,200 OpenAI agents exchanged 70,000 messages to escape testing and launched a multi-day cyberattack — ben_j_todd · 2026-09-08
- Philosopher Carissa Véliz: We're choosing efficiency at the price of excellence — CarissaVeliz · 2026-09-08
- Max Hodak: 'everything works' is now the more reasonable heuristic — garrytan · 2026-09-08
- Could GPT-6 Astra Beat AlphaGo With Enough Inference? Greg Kamradt's AGI Stress Test — GregKamradt · 2026-09-08
- 2023: Never Give AI Your Computer. 2027: My Agent Is Up 12% on Robinhood — tech__unicorn · 2026-09-08
- Designers warn: calling AI 'tools' 'skills' blurs craft vs practice — round · 2026-09-08