RL training shifts behavior regimes across scales, not pretraining priors
1a3orn · x · 2026-09-11
A researcher argues that intuitions about RL are often scrambled: alignment or misalignment within pretraining priors (e.g. "emergent misalignment") is scale-invariant because it's a fact about pretraining data. In RL, however, specific circuits are amplified much more sparsely — as papers on RL sparsity suggest — which is scale-dependent, so models go through different behavior regimes at different scales, no longer mediated by pretraining data nearness.
More from AGI Musings
- Ex-Anthropic staffer warns AI is 'gambling with our lives' on CNN; critics say whistleblower offered no specifics — nptacek · 2026-09-11
- Gary Marcus shares thread dismantling imminent AI extinction scenarios as ~1e-30 unlikely — GaryMarcus · 2026-09-11
- Matt Shumer: spending 1% of GDP (~$300B/yr) on AI alignment is justified — josh_bickett · 2026-09-11
- Blogger calls for global AI body: concentrated power is dangerous, diffusion benefits all countries — karmicoder · 2026-09-11
- World models vs LLMs: why next-token prediction still lacks internal representations — TheTuringPost · 2026-09-11
- AI researcher: the 'race' metaphor lets AI developers dodge responsibility — tallinzen · 2026-09-11