Dev finds 1x1 convs with constant weights best squeeze TFLOP/s from Pixel 9A's TPU
mgostIH · x · 2026-10-09
While debugging the Pixel 9A's G4 TPU compiler, mgostIH found matmul performance is highly sensitive to compiler preferences: 1x1 convolutions with constant parameters are the best way to squeeze out TFLOP/s, but you must find the compiler's happy path or compilation times out. First-hand on-device TPU tuning experience.
Related event: Pixel 9A TPU compiler slammed by developer, performance data hidden(3 posts)→
More from Infra
- Proximal Opens Up Its Internal Post-Training Stack for High-Quality Data — brianryhuang · 2026-10-09
- David Manheim: open models shift AI power not to users but to NVIDIA — davidmanheim · 2026-10-09
- When compute debt outpaces user cash: margins flow from software to physical asset owners — LexSokolin · 2026-10-09
- NVIDIA Dynamo adds session-aware inference: reuse agent KV cache across vLLM and SGLang — PyTorch · 2026-10-09
- Signal65: A model that fits in 128GB now matches Claude Opus 5 on agentic work, within 5% of a 2.4T flagship — ryanshrout · 2026-10-09
- Developer gets a full H100 node running at home — TheZachMueller · 2026-10-09