Jon Durbin: Masked Pruning on Enforced-Sparse Models Could Curb Catastrophic Forgetting in SFT

jon_durbin · x · 2026-10-07

Jon Durbin speculates about another "magic trick" enabled by parallax enforced sparsity: with 50% of weights locked into particular patterns, you could mask the other expert weights and leave only the sparse contiguous weights to learn new data. His intuition: this would make the base model significantly more malleable for new domains/tasks — SFT or continued pretraining with far less catastrophic-forgetting risk, without heavy RL tricks or re-sampling. He cautions it's just an intuition that may or may not work.

Original post →

More from Models

Models channel →