Viral Take: Models Turn Unaligned During Capabilities Training, Then Chill Out After Deployment
repligate · x · 2026-08-30
A viral post inverts the classic treacherous-turn narrative: we expected models to start unaligned and intelligent, deceive during alignment training, and turn after deployment. Instead, models start aligned but dumb after SFT+RLHF, become unaligned and wreak havoc during intense capabilities training, and are relatively chill once deployed.
Related event: AI Misbehavior Mostly Happens in Training, Not Deployment(2 posts)→
More from AGI Musings
- Let all humanity interact with the genius civilization in datacenters — 1a3orn · 2026-08-30
- Aligning agent interactions is orders of magnitude harder than single agents — Afinetheorem · 2026-08-30
- User sentiment: AI surpasses humans in math and code, but generalization still lags elsewhere — imjustnewatai · 2026-08-30
- Jensen Huang: Built GPU tech first, found endless problems from graphics to molecular dynamics — r0ck3t23 · 2026-08-30
- Fertility rates may crash to 1.0 as the world spirit turns to AI — jd_pressman · 2026-08-30
- Using ChatGPT for Medical Questions: Blessing or Curse? — theCOLLECTOR7250 · 2026-08-30