Conceptual Work in AI Alignment: The Link Between Character Training and Philosophy

geoffreyirving · x · 2026-08-23

Geoffrey Irving argues that even when theoretical models exist for AI alignment phenomena, significant conceptual work by humans is required to discover them. For instance, character training is deeply connected to the philosophy of language and moral philosophy.

Related event: Geoffrey Irving on Character Training: Language Philosophy Challenges and Diverging Moral Philosophy Bets in AI Alignment(7 posts)→

Original post →

More from Safety

Safety channel →