M1 Pro Benchmarks: Qwen 3.8 Hits 45 tok/s at 256k Context
EyalToledano · x · 2026-09-02
Eyal Toledano shared benchmarks for running Qwen 3.8-Flash-Next on an M1 Pro (64GB RAM). With Q4 quantization and 32k context, it achieves 38 tok/s burst and 51 tok/s aggregate throughput with 4 agents. Even at 256k context using Q3 quantization, it sustains 18 tok/s and reaches 57 tok/s aggregate throughput.
Related event: Qwen3.8-Flash-Next with pMLX Engine Runs Big Models on Small Memory(3 posts)→
More from Infra
- Intel exec: AI era security requires silicon-level design, not afterthoughts — BenBajarin · 2026-09-02
- PyTorch 2.14 released with 2,995 commits from 487 contributors — PyTorch · 2026-09-02
- Exllamav3 benchmarks: 700tk/s on 8x3090 setup — Leflakk · 2026-09-02
- GLM-5.3 Model Gets GGUF Quantization Release for Edge Deployment — unsloth · 2026-09-02
- Asus AI PC Price Jumps 50%, Speculating on Upcoming DGX Spark Hike — mountainyoo · 2026-09-02
- User Praises GPT Infra Stability: Months Without Downtime — natesiggard · 2026-09-02