Streaming MoE Experts from SSD: DeepSeek V4.1 Hits ~40 tps via mlx-stream
HankYeomans · x · 2026-10-08
A writeup on the mlx-stream extension for mlx-serve shows DeepSeek-V4.1-Flash streaming its MoE experts from SSD. Key optimization: each layer's router now runs before attention completes, so expert reads start early—at 16K tokens, 81% of reads are issued ahead of time—yielding nearly 40 tokens/s decode on Apple Silicon.
More from Infra
- ~90% of frontier lab compute now goes to post-training and inference: Social Capital — soumitrashukla9 · 2026-10-08
- Windows demo routes coding tasks to local model with GPU spinning, llama.cpp lands on Windows ML — ryanshrout · 2026-10-08
- exe.dev explains crossing the hyper-thread boundary: core scheduling cookies for VM isolation — davidcrawshaw · 2026-10-08
- NVIDIA's LoGRA cuts RL training memory by up to 45.7%, trains 27B model where Adam OOMs — mark_k · 2026-10-08
- Qwen3.8-Flash-Next on 6x3090 without NVLink: prefill 8-10x faster, long-context decode 2-3x — flynth92 · 2026-10-08
- Google DeepMind's Philipp Schmid: Give Every AI Agent Its Own Managed Cloud Sandbox — AI Engineer · 2026-10-08