Apple Silicon Inference Optimization Is a Mess: vllm-metal Closest to Complete Stack, mlx-lm Lacks Key Features

McFlurriez · reddit · 2026-08-16

After two weeks of research, the author finds the Apple Silicon inference optimization ecosystem fragmented, lacking integrated optimizations like CUDA. Key issues: Qwen's hybrid KV/recurrent state complicates prefix caching and speculative decoding; mlx-lm drops MTP heads during conversion. vllm-metal is closest to a complete stack; the author urges consolidating efforts.

Original post →

More from Infra

Infra channel →