Streaming MoE Experts from SSD: DeepSeek V4.1 Hits ~40 tps via mlx-stream

HankYeomans · x · 2026-10-08

A writeup on the mlx-stream extension for mlx-serve shows DeepSeek-V4.1-Flash streaming its MoE experts from SSD. Key optimization: each layer's router now runs before attention completes, so expert reads start early—at 16K tokens, 81% of reads are issued ahead of time—yielding nearly 40 tokens/s decode on Apple Silicon.

Original post →

More from Infra

Infra channel →