Sail Research doubles Gemma 4 31B prefill MFU to 63% on TPU v6e

xennygrimmato_ · x · 2026-09-23

Sail Research details its optimization journey serving Gemma 4 31B on Google's TPU v6e, lifting prefill MFU from 32% to 63%, and open-sourced HTDYM, its internal performance modeling tool for choosing the most cost-effective deployment chips.

The post breaks down v6e's odd specs: BF16 FLOPs matching an H100 but 2.5x less HBM, less than half the memory bandwidth, and a neighbor-only ICI topology (2x2 ring) with hard-to-compare official bandwidth figures — measured host-wide AllGather reaches 180 GB/s effective.

Key takeaway: depending on the model, niche accelerators can deliver big-player performance at a substantial discount, and Gemma 4 31B fits v6e well.

Original post →

More from Infra

Infra channel →