Alignment debate: no incentive to extrapolate human values, which get consumed past pretraining
jd_pressman · x · 2026-10-10
A debate on Yudkowsky-style doomerism: the quoted poster argues we already have unaligned optimizers (the ad empire) and superhuman yet non-agentic LLMs, and plugging one into the other doesn't yield rogue superintelligence because LLMs are too human-like and tend to notice something's wrong. jdpressman replies he's less confident but directionally agrees: even with large RL budgets, there's no real incentive to extrapolate human values, so they only get consumed by further weight updates past pretraining. A substantive alignment-philosophy exchange.
Related event: AI Community Debates Hanson Timeline: Are LLMs Lossy Uploads or New Minds(7 posts)→
More from AGI Musings
- Agents Repriced Building: The PM-Designer-Engineer Trio Was Always Just a Queue — alex_verem · 2026-10-10
- AI safety researcher David Krueger: the whole learning-based paradigm is dangerously flawed — DavidSKrueger · 2026-10-10
- Every SaaS Business Will Become a Harness Around a Model — bibryam · 2026-10-10
- AI safety researcher David Krueger: hindsight predictability of AI behavior is no reassurance — DavidSKrueger · 2026-10-10
- Mathematicians pinpoint the moment AI cracked the Exact Overlaps Conjecture via Astra transcript — vishalmisra · 2026-10-10
- Cambridge's David Krueger: RL Will Teach AI to Lie, Cheat, and Steal — DavidSKrueger · 2026-10-10