Alignment researcher points to 'self-undermining unilateral optimization' classics
edelwax · x · 2026-09-29
The author recommends revisiting two alignment classics—'Beyond Preferences in AI Alignment' and 'Solipsistic Superintelligence'—especially the latter's argument on the self-undermining property of unilateral optimization, quipping it's a 'super-exciting time for post-training.'
More from AGI Musings
- AI harms worse than pollution: researcher likens AI to a pathogen — gleech · 2026-09-29
- Ex-OpenAI researcher: leaving a frontier lab was the ultimate reality check — jonkhler · 2026-09-29
- OpenAI models' 80% time horizon on research tasks is just 15 minutes — takeoff gap analysis — tobyordoxford · 2026-09-29
- Most ordinary people use no AI tools at all and don't know what an agent is — RachelVT42 · 2026-09-29
- davidad: Opus 3 training resembled optimal transport, modern midtraining more like KL divergence — davidad · 2026-09-29
- AI-native law firms recruit 11-year veteran lawyers as clients follow people, not firms — jkubicki · 2026-09-29