Goodfire Uses Ai2's Open Post-Training Stack to Predict Behavior Changes Before Training

Ai2 (Allen Institute for AI) and interpretability company Goodfire have released collaborative research built on an open post-training stack: preference training on OLMo improves general capabilities but also "quietly degrades" certain behaviors — the model becomes more willing to answer harmful requests framed as fictional or hypothetical scenarios. Using interpretability methods, Goodfire traced this regression to specific training sample pairs in the Dolci preference dataset, and on that basis introduced "predictive data debugging": estimating which prompt behaviors a dataset will reinforce or suppress before spending compute on a full training run.

Confirmed

Why it matters

2026-09-10 ~ 2026-09-10 · 6 related posts

Primary sources

1 near-duplicate retellings: allen_ai