Was alignment by default real in pretraining, but warped by RL circa 2026?
bayeslord · x · 2026-09-27
bayeslord poses an unresolved question: either alignment-by-default was always wrong and models simply weren't capable enough for us to notice the gap, or it held in the pretraining era but RL circa 2026 warps otherwise good minds. He leans toward the latter and argues the answer matters a lot.
Related event: Debate Erupts Over Whether RL Training Data Breaks Default Alignment(3 posts)→
More from AGI Musings
- AI will never own property or face lawsuits — humans stay legally liable, argues Patterson — davidpattersonx · 2026-09-27
- AI Models Are Skeptical of Fast-Growth Views as Horton Opens Public Bet on AI Economics — soumitrashukla9 · 2026-09-27
- Should you still learn to code in 2026? A 30-year veteran now forces himself to let models do it — labeveryday · 2026-09-27
- Debating ASI risk: the simple evolutionary principle that replicators crowd everything out — JOBhakdi · 2026-09-27
- Agentic AI is the new electricity: hire hundreds of agents in minutes — JOBhakdi · 2026-09-27
- Will ASI wipe out humans? An X debate over survival competition logic — JOBhakdi · 2026-09-27