A Rough Character Training Story: Self-Generating Ethical Data
geoffreyirving · x · 2026-08-23
Geoffrey Irving outlines a rough character training story related to philosophy of language: 1. Train the model; 2. Tell it to be 'ethical'; 3. Ask it to generate 'ethical' data for itself; 4. Train further; 5. ...? This illustrates a methodology of reinforcing specific traits through self-generated data loops.
More from Safety
- Clarifying 'AI Security': Model Security vs AI for Security — EarlenceF · 2026-08-23
- DOJ reportedly probing a16z over board seats at competing startups — HaktanSuren · 2026-08-23
- 2026 International AI Safety Report: Limited Evidence of AI Manipulation, Positive Learning Impact — flowersslop · 2026-08-23
- AI Alignment Banter: User Challenges Zvi to Predict All LLM Misuse Cases — jd_pressman · 2026-08-23
- Developer Releases Tool to Remove SynthID Watermarks from AI Images — Great-Investigator30 · 2026-08-23
- AI Fairness Discourse Recycled into AI Safety Discourse — rajiinio · 2026-08-23