Qwen 3.8 Flash Next runs at 65 t/s on M3 Ultra, Q2 weights released on Hugging Face
ivanfioravanti · x · 2026-09-07
Developer ivanfioravanti shows Qwen 3.8 Flash Next running locally on Apple Silicon via DwarfStar, hitting 65 tokens/s on an M3 Ultra in a video demo.
He uploaded Q2 quantized weights to Hugging Face (ivanfioravanti/Qwen3.8-Flash-Next-DS4-IQ2) and opened a PR branch, looking for volunteers with 64GB Apple Silicon machines to test. The model's chat template exposes reasoning effort levels of xhigh (default), medium, and low.
More from Infra
- Report: AWS raises 2026 capex to $220B, ramps AI server rack orders via TSMC, Alchip, Wiwynn, Foxconn — firstadopter · 2026-09-07
- Together AI signs one of industry's largest open-source infra deals: 120,000 chips, ~$5B/yr — togethercompute · 2026-09-07
- Firecracker jailer security layers dissected: io_uring defeats a key sandbox boundary — jedisct1 · 2026-09-07
- Perplexity goes local: private tasks hand off to on-device small models — HowDevelop · 2026-09-07
- DeepSeek V4 Flash at 75% off via Merge Gateway: $0.04/M input tokens through Sept 30 — shensi · 2026-09-07
- Homelab With 4x RTX 4090 Weighs vLLM+P2P Patch vs llama.cpp for Qwen Models — dowitex · 2026-09-07