Running MiMo 2.6 Flash across an RTX 6000 and M5 laptop at 40 tokens/sec over 10 GbE

pcuenq · x · 2026-10-07

HF engineer pcuenq demos heterogeneous local inference: MiMo 2.6 Flash with native mxfp4 weights, distributed across an RTX 6000 Blackwell and an M5 laptop over 10 GbE, hitting 40 tokens/sec out of the box in llama.cpp. He argues local AI is closing in on frontier models, and you don't always need the biggest closed model.

Original post →

More from Infra

Infra channel →