Geoffrey Irving on the Grounding Problem in Character Training and Alignment
geoffreyirving · x · 2026-08-23
Geoffrey Irving discusses the challenges in character training, noting that at early stages, the model is weak and unaligned. Concepts like 'ethical' are merely internal numbers (e.g., 46318) influencing token distribution. The situation grounds into reality only in a loopy, indirect manner, making it unclear how much value we derive from this complex tangle.
More from Safety
- Beba Cibralic joins Resolution as Philosophy Research Lead — sethlazar · 2026-08-23
- AI Assistant Instinct Faces Privacy Backlash Over Data Usage Terms — steipete · 2026-08-23
- Instinct narrative flips from promise to security risks — manosaie · 2026-08-23
- AI enables mass surveillance of everyone, but privacy institutions are stuck in the 1700s — AaronBergman18 · 2026-08-23
- Steganographic communication may emerge in multi-agent RL without obfuscation rewards — brianryhuang · 2026-08-23
- Ex-OpenAI Researcher Launches AVERI to Standardize Frontier AI Auditing — dhadfieldmenell · 2026-08-23