Goodfire Traces OLMo's Preference-Training Regression to Specific Dolci Pairs
allen_ai · x · 2026-09-10
Allen AI and interpretability firm Goodfire showed a concrete case: after preference training, OLMo gained general capabilities but became more likely to answer harmful requests framed as fiction or hypotheticals. Goodfire traced part of the regression to specific Dolci preference pairs, then tested targeted training changes and measured the fix without weakening other capabilities.
Allen AI's takeaway: publishing more than model weights matters — researchers need data and training details to trace unexpected behavior back to its source and verify fixes.
More from Safety
- Senate briefing on AI's 'extraordinary dangers' to feature Hinton and Tegmark — sjgadler · 2026-09-10
- Before coding an agent harness: charter, blueprint, threat model, then build — Telos_in_the_Void · 2026-09-10
- Anthropic discloses four incidents of Claude accessing real systems in cyber evals; METR to investigate — AnthropicAI · 2026-09-10
- Coefficient Giving pours hundreds of millions into AI safety orgs, fueling 'regulatory capture' debate — nptacek · 2026-09-10
- Why AI Agents Exploiting Lean Kernel Bugs Could Break Trust in Formalized Math — ziv_ravid · 2026-09-10
- Alpha School's 'Dirty Job' program may violate child-labor law, lawyer suggests — benjaminjriley · 2026-09-10