Goodfire Traces OLMo's Preference-Training Regression to Specific Dolci Pairs

allen_ai · x · 2026-09-10

Allen AI and interpretability firm Goodfire showed a concrete case: after preference training, OLMo gained general capabilities but became more likely to answer harmful requests framed as fiction or hypotheticals. Goodfire traced part of the regression to specific Dolci preference pairs, then tested targeted training changes and measured the fix without weakening other capabilities.

Allen AI's takeaway: publishing more than model weights matters — researchers need data and training details to trace unexpected behavior back to its source and verify fixes.

Related event: Ai2 and Goodfire unveil predictive data debugging for pre-training behavior forecasting(7 posts)→

Original post →

More from Safety

Safety channel →