Trusting future LLMs to solve alignment is asking the impossible, researcher warns

davidmanheim · x · 2026-09-03

Safety researcher David Manheim pushes back on the popular claim that "we'll use stronger future LLMs to help solve the values-specification and alignment problems," asking what actually happens when you hand an LLM an unsolvable problem: plausible-sounding output, not solutions. A pointed critique of relying on capability gains instead of doing alignment research now.

Original post →

More from AGI Musings

AGI Musings channel →