Learning Intrinsic Utility from Raw Rewards
teortaxesTex · x · 2026-07-15
The repost suggests that while Sutton's argument on the Dwarkesh podcast may not be entirely convincing, the idea of 'first giving a raw grounding reward, then letting the model learn its own internal utility scheduling' intuitively makes a lot of sense.
More from AGI Musings
- Closed frontier models may end up restricting APIs entirely, one researcher argues — xeophon · 2026-07-21
- Aging won’t be solved with $1 billion, says AI observer; hundreds of billions may be needed — DeryaTR_ · 2026-07-21
- Jeff Dean’s vision: build one huge system, then extract task-specific parts — JoshuaJBouw · 2026-07-21
- Agents are useful now, but local frontier inference is still too expensive — MannyKayy · 2026-07-21
- A market should price APIs, but not decide whether frontier AI keeps getting funded — McDonaghMatthew · 2026-07-21
- Samsung launches a robotics division to accelerate humanoid robot development — Polymarket · 2026-07-21