M5 Ultra Shootout: Qwen3.8-Flash-Next Prefills 200K in 47s vs Laguna's 455s

nonlinearsystems · reddit · 2026-10-01

A same-day, same-harness benchmark on Mac Studio M5 Ultra 256GB: Qwen3.8-Flash-Next (oMLX, MTP) holds 4,200 tok/s linear prefill, hitting 200K context in 47.1s, while Laguna-S-2.1 degrades superlinearly to 455.4s. Decode: Qwen 59-74 tok/s with 70-76% MTP acceptance; Laguna falls from 68 to 34. Quality tied 4/4 on script-verified problems. Key pitfall: shared prefixes let KV cache leak across runs, faking 21s prefill.

Original post →

More from Infra

Infra channel →