Quintin Pope cites Ngo-Yudkowsky transcripts on why SGD learns dangerous search from safe domains

QuintinPope5 · x · 2026-10-08

Interpretability researcher Quintin Pope invoked MIRI's transcribed Ngo–Yudkowsky Discord dialogues on alignment difficulty (the series also includes Ajeya Cotra, Paul Christiano, Carl Shulman and others) in a debate. Key point: asked why SGD-trained narrow-capability models lead to general ones, Yudkowsky argued the optimization process inherently tends toward learning dangerous search even from 'safe' domains — which Pope cited while arguing Yudkowsky mispredicted capability entanglement within a single SGD-produced system.

Related event: Quintin Pope Debates Yudkowsky's Model of Intelligence and Generalization(7 posts)→

Original post →

More from AGI Musings

AGI Musings channel →