Specialized Apple Silicon stacks beat LM Studio by 2x with native MTP speculative decoding

AccBalanced · x · 2026-09-13

A reposted thread lists three cases where model-specific Apple Silicon inference stacks beat general-purpose ones: @antirez built DS4/DwarfStar specifically around DeepSeek V4 Flash instead of waiting for a universal runtime; @Youssofal's MTPLX brought native MTP speculative decoding to Apple Silicon before MLX/GGUF/LM Studio supported it; and @ddalcu's mlx-serve shows +154% from native MTP on Qwen 3.6 35B, roughly 2x vs LM Studio in benchmarks.

The argument: aggressively exploiting a model's architecture and a hardware target is becoming significantly more optimal than ecosystem-generic stacks, and this trend grows with LLMs.

Original post →

More from Infra

Infra channel →