CapTrack, accepted at NeurIPS, finds LLM forgetting is highly heterogeneous across capabilities after post-training
schwarzjn_ · x · 2026-09-29
CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training has been accepted to NeurIPS 2026 (arXiv:2603.06610).
- Redefines forgetting beyond parametric/factual knowledge loss as systematic model drift that degrades behavior and user experience
- Introduces a capability-centric framework combining a behavioral taxonomy with capability-specific evaluation metrics
- Large-scale empirical study across post-training algorithms, domains and model families, up to 80B parameters
- Key findings: forgetting is highly heterogeneous across capabilities, with pronounced drift in robustness and default behaviors; instruction fine-tuning induces the strongest relative drift, while preference optimization is more conservative and can partially recover lost capabilities; differences across model families persist, and no universal mitigation emerges
- Practical warning: if you touched the weights during post-training, check what changed across capabilities
More from Research
- AWS's Marc Brooker built the best 2B decision model — for about 24 hours — ShadajL · 2026-09-29
- JevBench Scales to 6x Test Cases, Rotates Held-Out Sets and Penalizes Benchmaxxing — airesearch12 · 2026-09-29
- Watch a Solo Dev Post-Train an 80B Model at Home on V100s: 96 Hours of Distillation, 3340 Samples — jjusko20 · 2026-09-29
- Sperm whales actively exchange vowels in dialogues, suggesting compositional codas — begusgasper · 2026-09-29
- NVIDIA ICRA'26 keynote: human data is the most scalable source for robot foundation models — yukez · 2026-09-29
- Prime Intellect brings multi-agent training to its open RL stack PRIME-RL — willcb · 2026-09-29