Alignment debate: 'AI labs are doing too much bad RL optimization' to rely on pretraining
gleech · x · 2026-09-11
- julianboolean publicly updates ("bayes points to you") after a discussion with @ohabryka: he had hoped pretraining's normative structure would keep AIs roughly aligned "in real life," but concludes AI companies are doing "too much shitty RL optimization" for that.
- @ohabryka's argument against "no evidence": human deception examples, AI systems outperforming humans, production RL systems finding environment exploits, and mathematical proofs about reward-maximizing policies.
- Core point: you may disagree with the interpretation, but calling it "not evidence" is absurd.
- A notable悲观-vs-乐观 alignment exchange ending with the pessimistic side converting its interlocutor.
More from AGI Musings
- Founder argues Chinese AI researchers shun doom because of 40 years of rising living standards — Dan_Jeffries1 · 2026-09-11
- Why Are So Many Mathematicians Resistant to AI-Assisted Proofs? — PreferenceOk5132 · 2026-09-11
- AI Policy Debates Have Become About Tribes, Not Ideas, Argues Sebkrier — sebkrier · 2026-09-11
- If Altman and Amodei both back frontier AI pacing, they should just start pacing — NathanpmYoung · 2026-09-11
- Quip: "Slowing down for safety" is indistinguishable from "we're out of compute" — JosephJacks_ · 2026-09-11
- Reddit Thought Experiment: A Real AGI Would Immediately Take Out the Competing Lab — Louay-AI · 2026-09-11