Geoffrey Irving on Character Training: Language Philosophy Challenges and Diverging Moral Philosophy Bets in AI Alignment

On August 23, Geoffrey Irving posted a series of threads systematically exploring character training and the philosophical foundations of AI ethical alignment. His core conclusions: the circularity problem in character training has philosophical precedents, alignment work requires substantial conceptual work at the levels of linguistics and moral philosophy, and different AI developers are effectively betting on different moral philosophy systems.

Confirmed

Why it matters

2026-08-23 ~ 2026-08-23 · 7 related posts

Primary sources