Goodfire uses Ai2's open post-training stack to predict how training runs change model behavior
allen_ai · x · 2026-09-10
Ai2 highlights that training LLMs on preferred and dispreferred answers improves them while quietly worsening other behaviors. Goodfire AI used Ai2's open post-training stack to predict how a full training run would change responses to different prompts — effectively forecasting side effects of preference training before running it. Details in their technical thread.
More from Models
- DeepSeek v4.1 flash preview spotted running at ~300 tok/s, vision still missing — kevinkern · 2026-09-10
- OpenAI-compatible endpoint adds prepaid spend caps and per-key daily limits — testingcatalog · 2026-09-10
- User dumps Astra for Sol, burns through weekly token limit in 5 hours — StewartalsopIII · 2026-09-10
- Claude user burns through 98% of $200 Pro weekly limit with just $46 of API-equivalent usage — Angaisb_ · 2026-09-10
- Reddit user slams Gemini and Claude free tiers for manipulative "dark patterns" — Cheap-Manner-6164 · 2026-09-10
- One prompt burned 88% of weekly quota: Astra user hits a sub-agent usage drain — NanoIsAMeme · 2026-09-10