mlx-dspark speeds up Qwen3.8-27B by up to 3x on Apple Silicon
A-Rahim · reddit · 2026-08-15
mlx-dspark, an MLX port of DeepSeek's DSpark speculative decoding, now supports Qwen3.8-27B, achieving lossless speedups on Apple Silicon. On an M4 Pro 48GB, 8-bit quantization averages 2.45x speedup (3.0x math, 2.38x code, 1.96x chat), with code runs hitting 3.18x, boosting throughput from 8.3 to 20.3 tok/s. 4-bit gets 1.74x at 25.3 tok/s. Notably, 8-bit with speculative decoding (20-27 tok/s) beats plain 4-bit (14.6 tok/s), offering better quality at higher speed. The project includes an OpenAI-compatible server and Anthropic Messages API, enabling local model use with Claude Code, plus a native Mac app.
More from coding & agent
- Cloudflare open-sources its AI productivity environment, Cloudflare OS — hichaelmart · 2026-08-15
- Real-world example: Multi-agent decision making for test optimization — iamrobotbear · 2026-08-15
- Designing Payment Authorization for AI Agents: Balancing Security and Autonomy — NoCalendar831 · 2026-08-15
- Ex-Meta Scientist: Agents should use the web like humans via pixels and clicks — DhruvBatra_ · 2026-08-15
- OpenAI Codex Error: 'gpt-5.6-sol' Model Possibly Deprecated — AKsnipebuster47 · 2026-08-15
- reBot Arm Control Stack Integrates Agentic AI, VLM, and LLM for Robotics — kamathsblog · 2026-08-15