Stepped MoE paper: one model scales from 1B to 4B parameters, beating dense counterparts by 2-5%

pmttyji · reddit · 2026-10-09

A new arXiv paper unifies elastic architectures with sparsely gated MoE routing, conditioning the backbone on both context and target efficiency specs. The resulting model can run at 1/2/3/4B parameters, outperforming dense counterparts by 2-5% on knowledge-intensive benchmarks, matching static versions, with dense-level latency. Shared parameters across capacity points cut on-device disk usage and let serving adapt to available DRAM and compute.

Related event: Apple Proposes Stepped MoE for Elastic 1-4B On-Device Models(2 posts)→

Original post →

More from Infra

Infra channel →