Qwen3.8 Flash hits 45 tok/s on M4 Max, matching Qwen3.8 27B speed on Apple Silicon
DerTomsn · reddit · 2026-09-06
- A Reddit user benchmarked Qwen3.8 Flash Next local inference via llm-bench.io: 45 tok/s on M4 Max and 25 tok/s on M2 Ultra.
- The author reports it delivers similar speed to Qwen3.8 27B on Apple Silicon, signaling better small-model usability on-device.
More from Infra
- Independent Sweep Puts DiffusionGemma Peak Throughput at 3k tok/s, 3x Paper's Claim — bodonoghue85 · 2026-09-06
- 推理时给预训练 LLM 加滑窗注意力,64K 上下文 KV 缓存仅 3.5MB — ahsaor8 · 2026-09-06
- A One-Stop Dashboard to Find the Cheapest GPUs for Running MiniMax H3 — CategoryFew5869 · 2026-09-06
- Villager sim game built locally on 16GB VRAM with Qwen3.8-27B Q3 quant at 75 tok/s — Fancy-Snow7 · 2026-09-06
- US golf courses use 30.5x more water than all data centers combined — MaxUnfried · 2026-09-06
- Book-length deep dive: virtual memory from first principles — page tables, TLBs, NUMA and performance — abhi9u · 2026-09-06