Qwen3.8-27B launches on Mac with 933 tok/s prefill on M5 Max
max_paperclips · x · 2026-08-15
Qwen3.8-27B is now supported on MLX-VLM and NativAI upon release. Initial benchmarks on M5 Max 128GB show a Prefill speed of 933 tok/s and a Decode speed of 33 tok/s. The model maintains coherence and reasonable performance up to a 256k context length.
More from Infra
- Open source closes the gap with closed labs: Quality gap now just months — togethercompute · 2026-08-15
- How Nuclear Will Unlock Energy Abundance with Valar Atomics Founder — No Priors · 2026-08-15
- Nvidia reportedly cuts OpenAI data center guarantee from $250B to below $120B — rohanpaul_ai · 2026-08-15
- DeepSeek V4 on Cloudflare Workers AI: First 1M Token Context Models — ritakozlov · 2026-08-15
- New metric suggests dense models benefit significantly at bs=1 — teortaxesTex · 2026-08-15
- Regulators speed connections for AI facilities, require infrastructure payment — pstAsiatech · 2026-08-15