M5 Ultra Hits 3740 tok/s Prefill on Qwen, Nearly Double Overnight
EAccelerate_42 · x · 2026-09-22
- @mweinbach reports Qwen 3.8 Flash Next hitting 3740 tok/s prefill and 149 tok/s batched decode on an M5 Ultra, nearly double the previous day's numbers.
- @EAccelerate42 notes the config isn't even half of what 2x DGX delivers even for prefill, sparking debate over on-device inference performance.
More from Infra
- 63% of Americans oppose data centers in their community, spanning both parties, pollster says — AlexTensor · 2026-09-22
- Toby Ord estimates $20M spent on AI's millennium prize result, $200M for solid data — tobyordoxford · 2026-09-22
- What engineers check before adding a new LLM provider — Rama_Surasani_ · 2026-09-22
- Dev reverse-engineers DLSS 5 neural rendering, reimplements it bit-exact in Vulkan at 7.8ms/1080p — bdsqlsz · 2026-09-22
- Running 2.78T-param Kimi K3 on a single CPU in 8 GB RAM, no framework — tom_doerr · 2026-09-22
- LinkedIn's GPU Pre-Ranker Fuses Graph Features for 50x Bigger Models at 120ms p99 — _reachsumit · 2026-09-22