You Can't Punish a Model Into Alignment - Ajeya Cotra on Dwarkesh
Dwarkesh Patel · youtube · 2026-09-15
Dwarkesh Patel releases an interview with alignment researcher Ajeya Cotra titled "You Can't Punish a Model Into Alignment," examining why penalty-based training fundamentally falls short for value alignment in frontier models.
More from Safety
- Jensen Huang Bets on Cybersecurity as AI's Next Growth Market as Anthropic Disrupts Russian Hackers — coinfanking · 2026-09-15
- GlossoGen: new framework studies when LLM agents develop languages humans can't understand — lucy3_li · 2026-09-15
- AI arms-race narrative misleads policy, says researcher: cooperation, not competition, is the only path — GregCook2011 · 2026-09-15
- Yglesias: Backing chip exports to China while citing 'but China' is just pro-Nvidia, not policy — deanwball · 2026-09-15
- UN Security Council to hold meeting on AI next week amid international concern — scmp_news · 2026-09-15
- Big Tech's AI slowdown pact: safety agreement or outright cartel? — The Verge AI · 2026-09-15