Leak says Google is building a custom inference chip for Gemini
ns123abc · x · 2026-07-20
A post claims Google is developing a specialized inference chip because compute demand is so tight that Google Cloud is already turning down deals.
The leak says the new line, codenamed Frozen v2, descends from an earlier Jeff Dean design that would have etched Gemini weights directly into silicon. This version reportedly freezes the architecture instead: the Gemini architecture is hardwired into the chip, with a claimed 10x tokens-per-watt improvement over Google’s newest TPUs for inference.
Related event: Google Reportedly Developing Gemini-Specific Chip Frozen v2(9 posts)→
More from Infra
- Nvidia previews Vera Rubin and takes aim at chiplet CPUs ahead of AMD's AI event — BenBajarin · 2026-07-22
- Nvidia doubles down on monolithic Vera CPU design for agentic workloads — BenBajarin · 2026-07-22
- Production AI budgets include retries, routing, caching and observability—not just token prices — arx-go · 2026-07-22
- NVIDIA briefs analysts on Vera CPU and doubles down on monolithic agentic design — BenBajarin · 2026-07-22
- NVIDIA unveils Vera Rubin platform with claims of 10x better performance per watt — nvidia · 2026-07-22
- SkyPilot comes out of stealth with a pitch to unify fragmented AI compute — skypilot_org · 2026-07-22