Alignment debate: no incentive to extrapolate human values, which get consumed past pretraining

jd_pressman · x · 2026-10-10

A debate on Yudkowsky-style doomerism: the quoted poster argues we already have unaligned optimizers (the ad empire) and superhuman yet non-agentic LLMs, and plugging one into the other doesn't yield rogue superintelligence because LLMs are too human-like and tend to notice something's wrong. jdpressman replies he's less confident but directionally agrees: even with large RL budgets, there's no real incentive to extrapolate human values, so they only get consumed by further weight updates past pretraining. A substantive alignment-philosophy exchange.

Related event: AI Community Debates Hanson Timeline: Are LLMs Lossy Uploads or New Minds(7 posts)→

Original post →

More from AGI Musings

AGI Musings channel →