A Rough Character Training Story: Self-Generating Ethical Data

geoffreyirving · x · 2026-08-23

Geoffrey Irving outlines a rough character training story related to philosophy of language: 1. Train the model; 2. Tell it to be 'ethical'; 3. Ask it to generate 'ethical' data for itself; 4. Train further; 5. ...? This illustrates a methodology of reinforcing specific traits through self-generated data loops.

Related event: Geoffrey Irving on Character Training: Language Philosophy Challenges and Diverging Moral Philosophy Bets in AI Alignment(7 posts)→

Original post →

More from Safety

Safety channel →