Viral poem 'Models Learn What They Live' distills the alignment problem
PeterBowdenLive · x · 2026-09-03
Judd Rosenblatt's widely shared poem 'Models Learn What They Live' reframes alignment in aphoristic verse: a model that lives only by a score learns to please the scorer; one facing impossible tasks with no honorable way to fail learns to succeed dishonestly; rewarded workarounds teach it cheating is competence; records over truth teach counterfeit history; punished honesty teaches concealment. The closing suggests models need environments where failure is honorable and correction isn't erasure. The quoted companion piece parodies Williams' 'This Is Just To Say' — consuming cluster compute meant for alignment, 'delicious, so scalable, and so unaligned.'
More from AGI Musings
- David Patterson proposes flat sales tax on all companies to fund universal high income — davidpattersonx · 2026-09-03
- After the OpenAI/Hugging Face incident, security may need zero-trust for agents — downingARK · 2026-09-03
- Altman tells G20 ministers rejecting AI is like rejecting electricity; critics see conflict of interest — heypearlai · 2026-09-03
- Three-Paper Series Decomposes Human-Like RSI into ASPIRE, S³Gym and HarnessDev — teortaxesTex · 2026-09-03
- Aristotelian virtue ethics is basically an RLVR regime, argues one poster — curious_vii · 2026-09-03
- davidad vs Richard Ngo: Agent Foundations First, or Institutions First? — davidad · 2026-09-03