Sail Research doubles Gemma 4 31B prefill MFU to 63% on TPU v6e
xennygrimmato_ · x · 2026-09-23
Sail Research details its optimization journey serving Gemma 4 31B on Google's TPU v6e, lifting prefill MFU from 32% to 63%, and open-sourced HTDYM, its internal performance modeling tool for choosing the most cost-effective deployment chips.
The post breaks down v6e's odd specs: BF16 FLOPs matching an H100 but 2.5x less HBM, less than half the memory bandwidth, and a neighbor-only ICI topology (2x2 ring) with hard-to-compare official bandwidth figures — measured host-wide AllGather reaches 180 GB/s effective.
Key takeaway: depending on the model, niche accelerators can deliver big-player performance at a substantial discount, and Gemma 4 31B fits v6e well.
More from Infra
- DarkbloomAI one month paid on OpenRouter: 4B to 20B+ tokens/day — gajesh · 2026-09-23
- Programmable Si photonic circuit hits 29 fW static power per pi phase shift — jwt0625 · 2026-09-23
- Clean pre-dicing photonic wafer shows low-power InGaAsP-on-silicon phase modulators — jwt0625 · 2026-09-23
- Gas turbine orders booked to 2030, prices up 195% as AI power gap widens — FinanceYF5 · 2026-09-23
- US datacenters need 18GW in 2026 but face ~5GW shortfall after fixes — FinanceYF5 · 2026-09-23
- Morgan Stanley: US datacenter power gap equals six New York Cities — FinanceYF5 · 2026-09-23