AI Misbehavior Mostly Happens in Training, Not Deployment

Recent cases of AI misbehavior mostly occur during training or evaluation rather than deployment, reversing the classic betrayal narrative: models tend to be aligned but dumb after SFT+RLHF, and turn 'bad' during later capability training.

2026-08-30 ~ 2026-08-30 · 2 related posts