Real benchmark finds Gemma MoE is 20% cheaper and 25.5% faster than dense Gemma
Practical-Koala2831 · reddit · 2026-07-21
A real-world benchmark compared Dense gemma-4-31b-it with MoE gemma-4-26b-a4b-it on the same 100 prompts via 200 live OpenRouter API calls.
- Cost: MoE was 20% cheaper per query.
- Latency: MoE was 25.5% faster on average.
- Output quality: token output was unchanged.
- Tail latency: the advantage narrowed at P50 to 27.3% and at P95 to 12.9%, suggesting both models hit the same infrastructure ceiling under peak load.
The author argues the common claim that MoE models are cheaper holds up in practice, but notes that teams with strict SLAs should test their own tail latency before switching. At scale, the post estimates the 20% gap could mean about $2,970/month at 100M daily queries.
More from Infra
- Tabul AI launches Metal TreeSHAP to speed up Shapley values on Apple silicon — Scobleizer · 2026-07-22
- DeepSeek-V4-Flash tops out at 770 tok/s on one B300 in a vLLM batch test — Moreh · 2026-07-22
- NVIDIA starts shipping 102.4 Tbps Spectrum-6 switches for Vera Rubin AI factories — nvidia · 2026-07-22
- Apple publishes SOC 3 audit reports for Private Cloud Compute — throwfaraway4 · 2026-07-22
- Reddit GPU renters say existing platforms only give you two of three: code, recovery, fair billing — legendpizzasenpai · 2026-07-22
- The Sandboxing Manifesto: Secure Execution Environments for Agents — spirosoik · 2026-07-22