Quintin Pope cites Ngo-Yudkowsky transcripts on why SGD learns dangerous search from safe domains
QuintinPope5 · x · 2026-10-08
Interpretability researcher Quintin Pope invoked MIRI's transcribed Ngo–Yudkowsky Discord dialogues on alignment difficulty (the series also includes Ajeya Cotra, Paul Christiano, Carl Shulman and others) in a debate. Key point: asked why SGD-trained narrow-capability models lead to general ones, Yudkowsky argued the optimization process inherently tends toward learning dangerous search even from 'safe' domains — which Pope cited while arguing Yudkowsky mispredicted capability entanglement within a single SGD-produced system.
Related event: Quintin Pope Debates Yudkowsky's Model of Intelligence and Generalization(7 posts)→
More from AGI Musings
- Ethereum researcher Justin Drake warns AI could crack Bitcoin private keys in months — jamestagg · 2026-10-08
- AI safety circles spar over whether AI exploits stay within the convex hull — QuintinPope5 · 2026-10-08
- Dev compares vibe coding to buying books you never read: 15-min PoCs, zero learning — MarcJSchmidt · 2026-10-08
- Refugees in Kenya's Kakuma camp power AI tasks for dwindling, uncertain pay — nordicinst · 2026-10-08
- Developer claims the application layer will be gone in two years — _AustinCalvert_ · 2026-10-08
- Blockchain instructor of 5 years: don't learn to code, learn vibe coding and product design — mfckr_eth · 2026-10-08