M5 Ultra 256GB early test: Mimo2.6-Flash hits 49 tok/s at 32K context
bakawolf123 · reddit · 2026-09-22
A Reddit user got an M5 Ultra (36/80, 256GB unified memory) and hand-patched mlx-vlm to run Mimo2.6-Flash from original weights, sharing early benchmarks.
- 32K prompt / 64 generated tokens: 1238 tok/s prompt processing, 49.1 tok/s generation average
- 64K prompt / 1024 tokens: 982 tok/s prompt, generation drops to 38.1 tok/s, 90s total
- MTP is enabled but doesn't seem to be working; verdict is "pretty meh, but usable"
- Also posted oMLX numbers for qwen3.8-flash-next fp8 as comparison: 42-44 tok/s output, 177-188GB memory
- No out-of-the-box support exists yet; the raw-weights patch is the author's own work
More from Infra
- Same GPU, Four Ways to Buy It: What Nebius Spot Pricing Reveals About Compute Economics — demian_ai · 2026-09-22
- Tata Elxsi's IRIS: edge filtering cuts cloud frames 70-80% for real-time factory safety AI — AWS ML Blog · 2026-09-22
- bitsandbytes2 slightly delayed: zero-config lazy compression plus dynamic expert swaps for near-infinite KV cache — Tim_Dettmers · 2026-09-22
- Iterative sensitivity probing: how bitsandbytes2 finds each layer's compression limit — Tim_Dettmers · 2026-09-22
- Tim Dettmers releases runtime dynamic compression framework, hits 1.5-2.0 bit at high quality — Tim_Dettmers · 2026-09-22
- PSA: Non-US Users Should Consider Local AI in Case Governments Ban LLMs — TheMoonMidas · 2026-09-22