M1 Pro Benchmarks: Qwen 3.8 Hits 45 tok/s at 256k Context

EyalToledano · x · 2026-09-02

Eyal Toledano shared benchmarks for running Qwen 3.8-Flash-Next on an M1 Pro (64GB RAM). With Q4 quantization and 32k context, it achieves 38 tok/s burst and 51 tok/s aggregate throughput with 4 agents. Even at 256k context using Q3 quantization, it sustains 18 tok/s and reaches 57 tok/s aggregate throughput.

Related event: Qwen3.8-Flash-Next with pMLX Engine Runs Big Models on Small Memory(3 posts)→

Original post →

More from Infra

Infra channel →