M5 Ultra Shootout: Qwen3.8-Flash-Next Prefills 200K in 47s vs Laguna's 455s
nonlinearsystems · reddit · 2026-10-01
A same-day, same-harness benchmark on Mac Studio M5 Ultra 256GB: Qwen3.8-Flash-Next (oMLX, MTP) holds 4,200 tok/s linear prefill, hitting 200K context in 47.1s, while Laguna-S-2.1 degrades superlinearly to 455.4s. Decode: Qwen 59-74 tok/s with 70-76% MTP acceptance; Laguna falls from 68 to 34. Quality tied 4/4 on script-verified problems. Key pitfall: shared prefixes let KV cache leak across runs, faking 21s prefill.
More from Infra
- Local 27B Face-off: Dirk-Qwen3.8 Beats Swift-1.5 on a 200-Question Personal Eval — norenEnmotalen · 2026-10-01
- Google VP: a fraction of campus philanthropy could fund compute cloud for all university researchers — jasondeanlee · 2026-10-01
- Local Qwen costs ~€0.12/hour in electricity — is a frontier API actually cheaper? — soyalemujica · 2026-10-01
- Cognition First to Run NVIDIA Vera Rubin on CoreWeave, ~4.8x Throughput vs GB200 — silasalberti · 2026-10-01
- Azure AI kicks off series: why content extraction matters more as GenAI models get stronger — adnan_hashmi · 2026-10-01
- Micron Crushes Estimates as Quarterly Revenue Nearly Quadruples to $54.2B — Polymarket · 2026-10-01