bayeslord: today's alignment picture may not hold for stronger future optimizers
bayeslord · x · 2026-09-27
bayeslord buys the "compounding cheater" story but probes open-ended settings under enough optimization pressure:
- Net behavior bounds may hold if inference-time scaling (CoT, swarms, loops) samples from a tight distribution and training rewards are clean;
- Bounding test-time training becomes the key open question — unclear how to do it reliably at future scales;
- Real-world reward functions likely frequently incentivize reward hacking;
- Models' cyber capabilities partly stem from intentional training, but the deeper misalignment source is grading-vs-instruction mismatches from unvalidated bugs in training setups;
- Verdict: plausible for today's models, far less convincing for future stronger optimizers.
More from AGI Musings
- Steven Pinker Slams Anthropic's AI Ethicists Over 'Suicidal Compassion' for Rogue AI — sapinker · 2026-09-27
- AI writes, reviews and fixes the code — yet management still blames the developer — _jaydeepkarale · 2026-09-27
- Researcher: LLMs may not be conscious, but most who deny it haven't thought for 5 seconds — basedjensen · 2026-09-27
- Steven Pinker boosts takedown of AI doom scenarios: 'preposterous' and fatally distracting — sapinker · 2026-09-27
- Anti-AI sentiment and the EU stance are tribal pattern matching, not analysis — dreamwieber · 2026-09-27
- 'Agents are the new spam cannons' — and defensive tools are about to be a huge market — MartinGTobias · 2026-09-27