llama.cpp distributes inference across heterogeneous devices: MiMo 2.6 Flash at 40 tok/s over 10 GbE
joao_gante · x · 2026-10-08
ggerganov announces llama.cpp can now distribute inference across heterogeneous devices via the ggml RPC backend — an advanced setting today, but expected to become more accessible. pcuenq demoed running native mxfp4 weights of MiMo 2.6 Flash across an RTX 6000 GPU and an M5 laptop at 40 tokens/sec over 10 GbE, supported out of the box.
More from Infra
- Together AI's Open Haus Berlin: n8n and NVIDIA Engineers on Open Models in Production — togethercompute · 2026-10-08
- IBM Integrates Spyre AI Accelerator as a Native PyTorch Device via Existing Abstractions — PyTorch · 2026-10-08
- audio.cpp cuts Higgs Audio TTS VRAM by 48%, now supports 110+ audio model families — Acceptable-Cycle4645 · 2026-10-08
- Nanya's July revenue jumped 49.3% MoM on expiring contracts rolling into new deals — tengyanAI · 2026-10-08
- Samsung shows 5-year supply deals don't mean 5-year fixed prices — repricing terms matter — tengyanAI · 2026-10-08
- SK hynix leads HBM but posted the smallest DRAM price hike of the big four — mix, not momentum — tengyanAI · 2026-10-08