Stepped MoE paper: one model scales from 1B to 4B parameters, beating dense counterparts by 2-5%
pmttyji · reddit · 2026-10-09
A new arXiv paper unifies elastic architectures with sparsely gated MoE routing, conditioning the backbone on both context and target efficiency specs. The resulting model can run at 1/2/3/4B parameters, outperforming dense counterparts by 2-5% on knowledge-intensive benchmarks, matching static versions, with dense-level latency. Shared parameters across capacity points cut on-device disk usage and let serving adapt to available DRAM and compute.
Related event: Apple Proposes Stepped MoE for Elastic 1-4B On-Device Models(2 posts)→
More from Infra
- GPUs already within 2x of brain efficiency, and still beat human workers on energy — MikePFrank · 2026-10-09
- Do AI agents still need Kubernetes? Berlin event says yes, with agent-on-K8s cases — Al_Grigor · 2026-10-09
- Cloud Backlogs Hit $1.69T, CoreWeave Posts $2.58B Quarter as Inference Becomes the Battleground — FinanceYF5 · 2026-10-09
- NVIDIA Is AI's Central Bank: A100 Paper Citations Still Beat H100+H200 Combined — FinanceYF5 · 2026-10-09
- Browser-Based Calculator Crunches the Real Power Cost of Self-Hosted LLMs vs Cloud APIs — paq85 · 2026-10-09
- PartyKit shuts down free hosted platform 2.5 years after Cloudflare acquisition — threepointone · 2026-10-09