Cacheon Miners Push MiniMax M3 to 2,337.9 tok/s, +32.2% Over SGLang — Modeled as ~61% More Profit per Chip-Hour

JosephJacks_ · x · 2026-09-16

Cacheon reports its competition miners pushed MiniMax M3 inference throughput to 2,337.9 tok/s, +32.2% over the SGLang baseline. Using SemiAnalysis InferenceX modeling: a B300 on vLLM earns $4.81 profit per chip-hour at $4.25/hour cost; the same +32.2% throughput at constant pricing lifts modeled profit to $7.73/hour (+61%) — more capacity sold without buying more chips. The profit figure is a scenario, not measured; the throughput gain was reproduced via Cacheon's testing.

Original post →

More from Infra

Infra channel →