Qwen 3.8 runs 151% faster on Apple Silicon via community challenge
gajesh · x · 2026-08-16
The MLX community launched a performance challenge for the Qwen 3.8 27B model on Apple Silicon. With community contributions, the model now runs 151% faster than the baseline (53.3 tok/s decode speed). The project explores acceleration limits for local LLMs using speculative decoding.
More from Infra
- Nvidia's massive financing turns CUDA software support into hardware collateral — aronchick · 2026-08-16
- Developers enter the $0 cost per task era as inference prices plummet — DynamicWebPaige · 2026-08-16
- Qwen3.8-27B on 16GB VRAM: KV Cache Quantization Cliff from q4_0 to q4_1 — Unnamed-3891 · 2026-08-16
- MusCoRe protocol cuts agent history tokens by 71.9%, targeting edge inference on Pi 5 — Parallel_News · 2026-08-16
- Developer investigates undocumented B200 instructions for potential speedups — SkyLi0n · 2026-08-16
- Beating cuBLAS by 4.7%: NVFP4 Kernels Hand-Built on GB300, 100% Claude-Generated — pranjalssh · 2026-08-16