Goodfire Traces Olmo's Harmful-Answer Regression to Specific Dolci Preference Pairs
allen_ai · x · 2026-09-10
In an Ai2 x Goodfire experiment, preference training improved Olmo's general capabilities but made it more likely to answer harmful requests framed as fiction or hypotheticals. Using interpretability tooling, Goodfire traced part of the regression to specific preference pairs in the Dolci dataset — a concrete demonstration of debugging model behavior down to individual training examples.
More from Research
- VDiff-Bench: 1,756-question benchmark shows frontier models fail at spot-the-difference — yixin_wan_ · 2026-09-10
- AutoResearchExam uses hidden test sets to study how AI agents do 24-hour research — AlexGDimakis · 2026-09-10
- Goodfire's predictive data debugging previews how LLM training will change model behavior — leland_mcinnes · 2026-09-10
- OpenAI's claimed Navier-Stokes breakthrough ignored by mainstream media — IgorCarron · 2026-09-10
- Michael Levin's new paper: a structured latent space of patterns for new forms of life and mind — danfaggella · 2026-09-10
- Do AI doomers really have a strong forecasting record? XPT study suggests otherwise — random_walker · 2026-09-10