Jon Durbin: Masked Pruning on Enforced-Sparse Models Could Curb Catastrophic Forgetting in SFT
jon_durbin · x · 2026-10-07
Jon Durbin speculates about another "magic trick" enabled by parallax enforced sparsity: with 50% of weights locked into particular patterns, you could mask the other expert weights and leave only the sparse contiguous weights to learn new data. His intuition: this would make the base model significantly more malleable for new domains/tasks — SFT or continued pretraining with far less catastrophic-forgetting risk, without heavy RL tricks or re-sampling. He cautions it's just an intuition that may or may not work.
More from Models
- Google's multimodal embeddinggemma-2 trends on Hugging Face — google · 2026-10-07
- Mistral CEO: Large 4 trained on our own compute, 'RL shows no sign of saturation' — sivareddyg · 2026-10-07
- Perplexity ships open-weights pplx-decider-v1.1-27b at half the cost of v1 — perplexity_ai · 2026-10-07
- Mistral launches Large 4: 1T-param multimodal model, 49B active, open weights in October — beffjezos · 2026-10-07
- Perplexity's open-weights pplx-decider-v1.1-27b tops Hugging Face Decision Index 0.3 — AravSrinivas · 2026-10-07
- AutoAWQ Author: Reproduce Bonsai 2-Class Ternary Model for ~$43k on One B300 Node in ~4 Weeks — airesearch12 · 2026-10-07