Goodfire Traces Olmo's Harmful-Answer Regression to Specific Dolci Preference Pairs

allen_ai · x · 2026-09-10

In an Ai2 x Goodfire experiment, preference training improved Olmo's general capabilities but made it more likely to answer harmful requests framed as fiction or hypotheticals. Using interpretability tooling, Goodfire traced part of the regression to specific preference pairs in the Dolci dataset — a concrete demonstration of debugging model behavior down to individual training examples.

Related event: Ai2 and Goodfire unveil predictive data debugging for pre-training behavior forecasting(7 posts)→

Original post →

More from Research

Research channel →