Geoffrey Irving on Character Training: Language Philosophy Challenges and Diverging Moral Philosophy Bets in AI Alignment
On August 23, Geoffrey Irving posted a series of threads systematically exploring character training and the philosophical foundations of AI ethical alignment. His core conclusions: the circularity problem in character training has philosophical precedents, alignment work requires substantial conceptual work at the levels of linguistics and moral philosophy, and different AI developers are effectively betting on different moral philosophy systems.
Confirmed
- Irving noted that early-trained models are weak in capability and poorly aligned; "ethics" exists inside the model merely as a number affecting the token distribution (e.g., 46318 in new GPTs), and the connection between linguistic feedback loops and reality is concerning
- He outlined a simplified character training pipeline: train the model for a while, tell it to "be ethical," have it generate its own "ethical" data, keep training, with later steps still to be determined
- Borrowing Quine's argument against logical positivism, he noted that human language likewise cannot anchor each sentence to experiments in a directed acyclic graph (DAG) fashion, yet the entire linguistic web can coherently be anchored to reality as a whole—offering a reference for understanding circular language in AI models
- He argued that even if theoretical models exist for alignment, humans would need extensive conceptual work to discover them, and character training is deeply tied to philosophy of language and moral philosophy
- He observed that different developers bet on different moral philosophy systems: Anthropic leans toward virtue ethics, while OpenAI is closer to deontology (both mixed with elements of other systems)
Why it matters
- Character training is one of today's mainstream alignment practices, and its philosophical legitimacy directly affects the credibility of reinforcement strategies built on self-generated data
- Different moral theories, extrapolated to superintelligence, may lead to vastly different outcomes; judging which theory is better bears on the value orientation of future advanced AI
2026-08-23 ~ 2026-08-23 · 7 related posts
Primary sources
- Conceptual Work in AI Alignment: The Link Between Character Training and Philosophy — geoffreyirving ·
- AI Developers Bet on Different Moral Philosophies: Anthropic on Virtue Ethics, OpenAI on Deontology — geoffreyirving ·
- Different moral theories shape superintelligence differently; Anthropic leans virtue ethics while OpenAI favors deontology — geoffreyirving ·
- [source] Conceptual Work in AI Alignment: The Link Between Character Training and Philosophy — geoffreyirving · 2026-08-23
- A Rough Character Training Story: Self-Generating Ethical Data — geoffreyirving · 2026-08-23
- Geoffrey Irving on the Grounding Problem in Character Training and Alignment — geoffreyirving · 2026-08-23
- Drawing Parallels to Quine: The Loopy Nature of Human and AI Language — geoffreyirving · 2026-08-23
- Quine's Argument: The Entire Web of Language Grounds into Experiment Coherently — geoffreyirving · 2026-08-23
- [source] AI Developers Bet on Different Moral Philosophies: Anthropic on Virtue Ethics, OpenAI on Deontology — geoffreyirving · 2026-08-23
- [source] Different moral theories shape superintelligence differently; Anthropic leans virtue ethics while OpenAI favors deontology — geoffreyirving · 2026-08-23