MLX Community Makes Qwen 3.8 Flash Nearly 2x Faster on Apple Silicon, License Blocks Launch
gajesh · x · 2026-09-18
- The MLX.fast community has made Qwen 3.8 Flash nearly 2x faster on Apple Silicon, and plans to deploy it on Darkbloom, an open network of local Macs offering inference.
- The launch is on hold because the model's license requires a separate agreement for commercial serving. The author is publicly seeking a partnership with the Qwen team at Alibaba to make it happen.
More from Infra
- Google Open-Sources Agent Substrate on GKE: 10x Density, 1,000+ Dormant Agents per Host — blaizedsouza · 2026-09-18
- Redditor crams six V100 GPUs into a standard full-tower case for local LLM inference — Odd_Caterpillar_2994 · 2026-09-18
- Crusoe raises $3.9B at $30.9B valuation to build data centers and modular AI factories — TechCrunch AI · 2026-09-18
- A 2.5-hour first-principles primer on the semiconductor supply chain worth your time — blaizedsouza · 2026-09-18
- YC F26's Dreamscale Labs Moves Robot AI Inference to the Cloud — ycombinator · 2026-09-18
- Community fine-tunes an MTP head for Bonsai 2 27B, ~1.25x inference speedup — cephaloform · 2026-09-18