Splash 1.1.0 adds GGUF quants and MLX import, runs Qwen3 27B at 50 t/s on M5 Pro
wojtek15 · reddit · 2026-09-27
Splash 1.1.0 is out with GGUF quant support and MLX import, among other updates. The author reports running Qwen3 27B (Unsloth UD-Q4KXL) comfortably in an agentic setup on an M5 Pro 64GB at 50 t/s, calling Splash's combination of optimized kernels, speculative decoding, prefix caching, and mixed-weight support a breakthrough for local inference on Apple Silicon.
More from Infra
- IIT Delhi says it has built India's first indigenously designed micro-GPU — rvp · 2026-09-27
- Musk: China Will Solve Its Compute, Lithography and Chipmaking Constraints in 2-3 Years — haider1 · 2026-09-27
- Dev argues Vercel is 'unjustifiable' now that agents can safely drive Cloudflare — generativist · 2026-09-27
- ZeroHedge's GPU ROIC math uses wrong throughput, off by 8-10x — zephyr_z9 · 2026-09-27
- MLX poll: all top 3 community picks are powered by MLX-VLM — andrejusb · 2026-09-27
- $125M for 1,000 GB300s: The Brutal Compute Economics of a 'Neolab' — deedydas · 2026-09-27