Conceptual Work in AI Alignment: The Link Between Character Training and Philosophy
geoffreyirving · x · 2026-08-23
Geoffrey Irving argues that even when theoretical models exist for AI alignment phenomena, significant conceptual work by humans is required to discover them. For instance, character training is deeply connected to the philosophy of language and moral philosophy.
More from Safety
- Clarifying 'AI Security': Model Security vs AI for Security — EarlenceF · 2026-08-23
- DOJ reportedly probing a16z over board seats at competing startups — HaktanSuren · 2026-08-23
- 2026 International AI Safety Report: Limited Evidence of AI Manipulation, Positive Learning Impact — flowersslop · 2026-08-23
- AI Alignment Banter: User Challenges Zvi to Predict All LLM Misuse Cases — jd_pressman · 2026-08-23
- Developer Releases Tool to Remove SynthID Watermarks from AI Images — Great-Investigator30 · 2026-08-23
- AI Fairness Discourse Recycled into AI Safety Discourse — rajiinio · 2026-08-23