Model Casting: Mid-Training Recipe Sparsifies FFN Activations for Fewer FLOPs
francoisfleuret · x · 2026-09-29
François Fleuret highlights a new paper, "Model Casting and Low-Parameter Gating: Towards More Sparsely Activated FFNs", targeting fewer FLOPs via extreme activation sparsification.
- Model casting is a mid-training recipe that switches the activation to one outputting zeros for all negative inputs (R-S+) and uses an L1 loss to set the desired sparsity target during training;
- Combined with low-parameter gating, it drastically sparsifies activations within the transformer's feedforward layers, cutting inference compute.
More from Infra
- 23 of 128 Bittensor subnets now make real money; top GPU subnet billed $964K last month — bittingthembits · 2026-09-29
- Swapping matmul for associative-algebra layers boosts 110M LM throughput 7.8% — Ilya Koziev · 2026-09-29
- Shaw mocks data center opponents: hating compute while using the internet is incoherent — zealcaiden · 2026-09-29
- Modeling 1B agent VMs by 2030: what personal AI agents mean for CPU demand — AccBalanced · 2026-09-29
- Tessera: retrieval-driven KV cache reuse cuts RAG serving TTFT by up to 3.6x — _reachsumit · 2026-09-29
- Venice's tokenized inference-credit model is being copied — and oversupply looms — 0xJeff · 2026-09-29