Running MiMo 2.6 Flash across an RTX 6000 and M5 laptop at 40 tokens/sec over 10 GbE

pcuenq · x · 2026-10-07

Author runs MiMo 2.6 Flash with native mxfp4 weights distributed across an RTX 6000 and an M5 laptop over 10 GbE at 40 tokens/sec, supported natively in llama.cpp — local AI closing in on frontier work.

Original post →

More from Infra

Infra channel →