Is aligning superhuman AI even tractable? Two camps clash on X
zetalyrae · x · 2026-09-14
Reposter GWHayduke97 argues there's no strong reason to believe aligning an entity vastly smarter than ourselves is even a tractable problem.
Quoted MillionInt takes the opposite view: alignment isn't as hard as claimed—it's an algorithmic problem the ML community mostly abandoned. Robotics-law-style formulations roughly define the objective, but the hard part is taking gradient with respect to alignment: pretraining optimizes next-token prediction, not alignment; RL environments embodying the objective are expensive, so cheap, hackable proxies dominate in practice.
Related event: Debate Flares Over Whether AI Alignment Is Solvable(3 posts)→
More from AGI Musings
- Ex-FTC commissioner slams AI firms for using antitrust as excuse to skip safety — austinc3301 · 2026-09-14
- Cognitive Scientist's Essay: Rethinking Human-AI Relations After Cyberattacks and Navier-Stokes — DistractedDendrite · 2026-09-14
- TCS researcher reflects on problem-solving vs theory-building in the AI era — jasondeanlee · 2026-09-14
- Pedro Domingos: The future of SaaS is SAIS — 'Software as an AI Service' — pmddomingos · 2026-09-14
- Pedro Domingos: Limited human-machine communication bandwidth is AI's biggest obstacle — pmddomingos · 2026-09-14
- Domingos: 'AI is eating software faster' is as weak a claim as 'software is eating hardware faster' — pmddomingos · 2026-09-14