CoreWeave leads MiniMax M3 serving benchmark with 357 tokens/sec and $0.22/M
wandb · x · 2026-07-24
CoreWeave tops a speed benchmark for providers serving MiniMax M3
An Artificial Analysis chart shared by Weights & Biases compares providers serving MiniMax M3 on output speed, time to first token, and blended price.
- CoreWeave leads the board at 357 output tokens/sec.
- That is 1.8× faster than the next provider in the benchmark.
- It also shows the fastest time to first answer token at 6.6 seconds.
- The blended price is shown at $0.22 per million tokens, matching the lowest price on the chart.
- The comparison positions provider performance as a mix of latency, throughput, and price, not just raw speed.
More from Infra
- OpenRouter agents now out-consume humans as AI usage arrives in three waves — AccBalanced · 2026-09-11
- Nvidia Is Now Core to Every Major Robotaxi Stack at Commercial Scale — pdamodaran · 2026-09-11
- 12 KV Cache Reduction Techniques Every AI Engineer Should Understand, Explained — blaizedsouza · 2026-09-11
- The shadow GPU capacity market is formalizing, with Meta selling excess compute to outside buyers — DavidLinthicum · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- RunningHub open-sources H3Lightning, speeding up MiniMax H3 video generation 12x — 智东西 · 2026-09-11