Cacheon Miners Push MiniMax M3 to 2,337.9 tok/s, +32.2% Over SGLang — Modeled as ~61% More Profit per Chip-Hour
JosephJacks_ · x · 2026-09-16
Cacheon reports its competition miners pushed MiniMax M3 inference throughput to 2,337.9 tok/s, +32.2% over the SGLang baseline. Using SemiAnalysis InferenceX modeling: a B300 on vLLM earns $4.81 profit per chip-hour at $4.25/hour cost; the same +32.2% throughput at constant pricing lifts modeled profit to $7.73/hour (+61%) — more capacity sold without buying more chips. The profit figure is a scenario, not measured; the throughput gain was reproduced via Cacheon's testing.
More from Infra
- GPT-6 Astra reportedly uses loop transformers, adding compute depth instead of parameters — pranavmarla · 2026-09-16
- MediaTek's 2nm Dimensity 9600Pro runs 30B MoE models on-device with rebuilt dual NPUs — 量子位 · 2026-09-16
- Transformers models now run natively in vLLM with no port required — pcuenq · 2026-09-16
- TSMC builds 20 fabs yet can't meet AI demand as labor shortage slows expansion — emmanuelvivier · 2026-09-16
- SemiAnalysis: Nvidia Vera Rubin NVL72 Delivers Up to 30x Higher Throughput per MW Than Blackwell for Agentic Inference — emmanuelvivier · 2026-09-16
- Paying three AI vendors to parse our own docs: a Reddit quest for one self-hosted stack — Sad-Razzmatazz-7657 · 2026-09-16