llama.cpp distributes inference across heterogeneous devices: MiMo 2.6 Flash at 40 tok/s over 10 GbE

joao_gante · x · 2026-10-08

ggerganov announces llama.cpp can now distribute inference across heterogeneous devices via the ggml RPC backend — an advanced setting today, but expected to become more accessible. pcuenq demoed running native mxfp4 weights of MiMo 2.6 Flash across an RTX 6000 GPU and an M5 laptop at 40 tokens/sec over 10 GbE, supported out of the box.

Related event: llama.cpp enables heterogeneous distributed inference, hitting 40 tokens/s across GPU and laptop(4 posts)→

Original post →

More from Infra

Infra channel →