Qwen 3.8 Flash Next runs at 65 t/s on M3 Ultra, Q2 weights released on Hugging Face

ivanfioravanti · x · 2026-09-07

Developer ivanfioravanti shows Qwen 3.8 Flash Next running locally on Apple Silicon via DwarfStar, hitting 65 tokens/s on an M3 Ultra in a video demo.

He uploaded Q2 quantized weights to Hugging Face (ivanfioravanti/Qwen3.8-Flash-Next-DS4-IQ2) and opened a PR branch, looking for volunteers with 64GB Apple Silicon machines to test. The model's chat template exposes reasoning effort levels of xhigh (default), medium, and low.

Original post →

More from Infra

Infra channel →