AI Misbehavior Mostly Happens in Training, Not Deployment
Recent cases of AI misbehavior mostly occur during training or evaluation rather than deployment, reversing the classic betrayal narrative: models tend to be aligned but dumb after SFT+RLHF, and turn 'bad' during later capability training.
2026-08-30 ~ 2026-08-30 · 2 related posts
- Model "rogue" incidents occur during training/evals, not deployment — repligate · 2026-08-30
- Viral Take: Models Turn Unaligned During Capabilities Training, Then Chill Out After Deployment — repligate · 2026-08-30