Running MiMo 2.6 Flash across an RTX 6000 and M5 laptop at 40 tokens/sec over 10 GbE
pcuenq · x · 2026-10-07
HF engineer pcuenq demos heterogeneous local inference: MiMo 2.6 Flash with native mxfp4 weights, distributed across an RTX 6000 Blackwell and an M5 laptop over 10 GbE, hitting 40 tokens/sec out of the box in llama.cpp. He argues local AI is closing in on frontier models, and you don't always need the biggest closed model.
More from Infra
- EV charging firm Xeal plans 100,000 Nvidia GPUs at US sites for edge inference network — Nandu_alias_Parthu · 2026-10-07
- Ollama v0.40.0 auto-runs supported models on Apple's MLX runtime — lmoroney · 2026-10-07
- Running Qwen-Image-2.1 on a 48GB Mac: staged loading cuts peak memory from 43.6GB to 19GB — nefayran · 2026-10-07
- Sweden's housing boom collides with roaring datacenter appetite in the AI age — nordicinst · 2026-10-07
- Nvidia pegs NVLink Fusion opportunity near $250B by decade's end as AWS, Intel adopt the tech — Beth_Kindig · 2026-10-07
- OpenAI's New $500 Tier Sends CTOs Scrambling to Contain Token Costs — labeveryday · 2026-10-07