Google Reportedly Developing Gemini-Specific Chip Frozen v2

According to a report by The Information, Google is developing a dedicated inference chip codenamed Frozen v2 to alleviate severe compute shortages and serve Gemini more efficiently. The news remains a rumor, with no official confirmation from Google. The leak suggests that Google Cloud has even started turning down some orders due to extreme compute constraints.

Key Details

Frozen v2 is not a simple replacement for general-purpose TPUs. Its core approach is to hardcode parts of Gemini's model logic and architecture directly into the chip's circuitry to reduce computation and data movement during inference. According to @ns123abc and @markk, the team expects the new chip to achieve 6 to 10 times the tokens per watt (tokens/W) compared to Google's latest generation of TPUs. The chip is expected to launch as early as 2028. @ns123abc also noted that this chip line stems from a more aggressive design direction previously proposed by Jeff Dean.

Background and Impact

@Afinetheorem pointed out that this reflects a clear trend in the AI industry: companies are leaning towards "burning" more model logic into chips, significantly boosting efficiency through deep co-optimization of software and custom hardware. This customized approach combining hardware and software could alter the current market structure and competitive landscape.

2026-07-20 ~ 2026-07-22 · 10 related posts

Primary sources

7 near-duplicate retellings: ns123abc · ns123abc · toptickcrypto · mark_k · atShruti · 新智元 · ocean_protocol