Trusting future LLMs to solve alignment is asking the impossible, researcher warns
davidmanheim · x · 2026-09-03
Safety researcher David Manheim pushes back on the popular claim that "we'll use stronger future LLMs to help solve the values-specification and alignment problems," asking what actually happens when you hand an LLM an unsolvable problem: plausible-sounding output, not solutions. A pointed critique of relying on capability gains instead of doing alignment research now.
More from AGI Musings
- Nature Machine Intelligence paper proposes life-inspired interoceptive AI for adaptive agents — AnnaCiaunica · 2026-09-03
- Sequoia: the most valuable startups of the past 20 years rarely matched the hype — FinanceYF5 · 2026-09-03
- Philosopher gleech defends Bostrom's contested ASI paper: rhetoric damaged but argument intact — gleech · 2026-09-03
- Tech bubble underestimates how much ordinary people hate AI, warns SF founder — beffjezos · 2026-09-03
- Nine bold predictions: humanoid robots on market within 12 months, laptops beating frontier LLMs — StrategicHarmony · 2026-09-03
- Mihaela van der Schaar to open ECML PKDD2026 with AI discovery keynote — MihaelaVDS · 2026-09-03