Geoffrey Irving on the Grounding Problem in Character Training and Alignment

geoffreyirving · x · 2026-08-23

Geoffrey Irving discusses the challenges in character training, noting that at early stages, the model is weak and unaligned. Concepts like 'ethical' are merely internal numbers (e.g., 46318) influencing token distribution. The situation grounds into reality only in a loopy, indirect manner, making it unclear how much value we derive from this complex tangle.

Related event: Geoffrey Irving on Character Training: Language Philosophy Challenges and Diverging Moral Philosophy Bets in AI Alignment(7 posts)→

Original post →

More from Safety

Safety channel →