Goodfire uses Ai2's open post-training stack to predict how training runs change model behavior

allen_ai · x · 2026-09-10

Ai2 highlights that training LLMs on preferred and dispreferred answers improves them while quietly worsening other behaviors. Goodfire AI used Ai2's open post-training stack to predict how a full training run would change responses to different prompts — effectively forecasting side effects of preference training before running it. Details in their technical thread.

Related event: Goodfire's Predictive Data Debugging Forecasts Behavior Changes Before Training(6 posts)→

Original post →

More from Models

Models channel →