Splash 1.1.0 adds GGUF quants and MLX import, runs Qwen3 27B at 50 t/s on M5 Pro

wojtek15 · reddit · 2026-09-27

Splash 1.1.0 is out with GGUF quant support and MLX import, among other updates. The author reports running Qwen3 27B (Unsloth UD-Q4KXL) comfortably in an agentic setup on an M5 Pro 64GB at 50 t/s, calling Splash's combination of optimized kernels, speculative decoding, prefix caching, and mixed-weight support a breakthrough for local inference on Apple Silicon.

Original post →

More from Infra

Infra channel →