The design trades a general-purpose chip's flexibility for far cheaper AI inference

According to a report from The Information, Google is developing a new server chip, known inside the company as Frozen v2, that wires part of its Gemini model's architecture straight into the silicon.

Rather than replacing the general-purpose TPUs that run across many of Google's AI models, the chip is meant to be a lean, single-purpose offshoot of the company's chip family, built to run inference far more efficiently.

Internal engineering estimates suggest it could serve six to ten times more tokens per unit of energy than Google's newest TPUs. Baking part of the model's design into the hardware cuts the number of computing steps the chip has to perform and shortens the distance data must travel to produce a response, so answers come back faster.

The effort is aimed in part at easing a severe internal shortage of AI inference capacity — a squeeze that has stirred tension within Google and led Google Cloud to turn away some outside customers.

Frozen v2 descends from an earlier Frozen project credited to Google DeepMind chief scientist Jeff Dean, which would have embedded a model's full weights directly into the chip. That approach would have tied the hardware to a single model version. The new design is a compromise: it permanently fixes the underlying architecture while still allowing fresh weights to be loaded.

The work is early: how much of the architecture will ultimately be hardwired is undecided, and Google plans only limited trial production for internal use — at volumes expected to fall well short of its TPU output, with deployment targeted no earlier than 2028.

In a statement, Google said its teams are constantly exploring and testing new advances to give users and customers the strongest performance and efficiency, and that not every effort reaches production. Alphabet shares climbed by as much as 3.7% on Monday after the report.

Google is treating it as an exploratory bet rather than a product line, not a full-scale rollout.