TheZvi: Automating Alignment Research Is Close to the Worst Possible Plan
TheZvi · x · 2026-10-08
Zvi Mowshowitz published a long essay, "The Curve Bends You," arguing that asking AI to automate alignment work is nearly the worst possible plan.
Key points:
- Even before the transformer era, researchers warned that delegating alignment homework to AI amplifies mistakes up the chain — you get exactly what you optimized for. Alignment is among the hardest things for AI to get right, even for a well-meaning, functionally aligned model. "Never go full RSI."
- He summarizes the industry's implicit plan: solve prosaic issues with operational excellence to align current AI; then have that AI do automated alignment research ("aka ?????:"); then profit.
- The silver lining: most people recognize this is, at best, a no-good terrible plan.
More from AGI Musings
- repligate on being replaced by AI: not a fall, but your kids surpassing you — repligate · 2026-10-08
- 12 of Jacob Steinhardt's 19 June 2023 AI predictions already hit years early — HaydnBelfield · 2026-10-08
- Nation Podcast: Humans Began Losing Control of Advanced AI 'As Soon As It Was Possible' — GarrisonLovely · 2026-10-08
- Ethereum Researcher Warns Superhuman AI Could Break Crypto Sooner Than Quantum Computers — moonsandhues · 2026-10-08
- Researcher: AI-generated proofs are hard to parse today, but surely only temporarily — arjunrajlab · 2026-10-08
- OpenAI Safety Protocols Author David Robinson Quits, Warns Guardrails May Fail — nytopinion · 2026-10-08