Key Takeaways
- Frozen v2 is an Alphabet server chip that hardwires part of Gemini’s model architecture directly into silicon, rather than running Gemini on general-purpose hardware.
- Google engineers project six to ten times more tokens per unit of power than the company’s newest Tensor Processing Units. That figure is an internal projection, not an audited benchmark.
- The reported deployment target is 2028. Google has neither confirmed nor denied the project, first reported by The Information on 20 July 2026.
- The chip embeds the architecture, not the weights. New Gemini versions can be loaded onto it as long as Google keeps the same structural blueprint.
- An earlier version of the idea, credited to Google DeepMind chief scientist Jeff Dean, proposed burning trained weights into the chip. It was dropped because the hardware would age out with each model release.
- Frozen v2 sits alongside the TPU line rather than replacing it, and is unlikely to be sold to Google Cloud customers because it cannot run anyone else’s models.
- The motive is a compute shortage: Google Cloud’s backlog reached roughly $462 billion in Q1 2026, and Google rented Nvidia capacity from SpaceX for about $920 million per month as a stopgap.
- The main risk is architectural: if Gemini ever moves away from its transformer foundation, the hardwired portions lose their purpose.
Frozen v2 will power Gemini by removing work the hardware currently repeats billions of times a day. Instead of a general-purpose accelerator that recalculates Gemini’s structural requirements on every single request, the chip physically encodes parts of that structure into its circuitry. Fewer scheduling decisions get made at runtime, fewer intermediate values travel to memory and back, and each answer costs less electricity to produce. Google engineers put the gain at six to ten times the tokens per watt of the company’s newest TPUs, with deployment targeted for 2028.
What it will not do is make Gemini smarter. This is an inference chip, aimed at the serving side of the business rather than training. The practical payoff is capacity and margin: the same data centres, the same power contracts, the same grid connections, but far more Gemini responses per megawatt. For a company that told investors in April that it could not meet cloud demand, that arithmetic matters more than another benchmark win.
What Frozen v2 Actually Is
Frozen v2 is an application-specific integrated circuit built for one job: running Gemini-family models during inference. The name borrows from machine learning practice, where “freezing” a parameter locks its value so it stops changing. Here, a portion of the model gets locked into the hardware itself.
Google’s TPUs sit at the other end of the spectrum. They are custom accelerators, but they are still general-purpose within the AI domain — load any model onto them and they will run it. That flexibility is precisely what Google sells. It is also what costs energy. A chip that does not know in advance which model it will execute has to work out its own routing, scheduling, and memory access pattern every time a request arrives.
The reporting from The Information, picked up by CNBC and TechCrunch, described Frozen v2 as a deliberate narrowing of that scope. Google did not confirm it, saying only that its teams are “constantly researching and experimenting with new innovations” and that not every exploratory project reaches production.
Why General-Purpose Silicon Wastes Power on Gemini
Modern inference is less a mathematics problem than a logistics problem. The arithmetic itself is cheap. Moving data is not. Fetching model data from off-chip memory consumes roughly two orders of magnitude more energy than the arithmetic operations that data feeds. A model with billions of parameters cannot fit in a chip’s local cache, so the hardware shuttles weights and activations back and forth continuously. Adding faster chips does not fix that. Adding more chips does not fix it either. The only fix is to move less data.
Frozen v2 attacks that on two fronts. First, sequences of operations that a conventional chip executes as separate memory-touching steps can be fused into a single hardware primitive — the chip performs the whole compound operation in one pass and writes only the final result. Operator fusion already exists in software on standard accelerators; here it becomes permanent structure. Second, and more ambitiously, a chip designed entirely around one model’s shape could carry enough on-chip memory to host that model outright, which would come close to eliminating off-chip traffic during inference.
This is the same principle behind every specialist accelerator, including the tensor processing units Google pioneered. Frozen v2 simply pushes the specialisation one level further, from a class of workloads down to a single model family.
Architecture Versus Weights: The Distinction That Saved the Project
Frozen v2 exists because Frozen v1 did not work as a business case.
The original concept, credited to Google DeepMind chief scientist Jeff Dean, called for burning Gemini’s trained weights — the billions of numbers that determine how the model responds — straight into the circuitry. The efficiency would have been extraordinary. So would the obsolescence. A chip tied to one training snapshot has a useful life measured in months, and the engineering cost of a custom ASIC cannot be recovered on that timetable. The approach was set aside.
Version two separates two things that are easy to conflate. A neural network’s architecture is its blueprint: the number of transformer blocks, the count of attention heads, hidden dimension sizes, how feed-forward layers connect, which normalisation scheme is used. Those choices change rarely, and only when a lab rethinks its design. The weights are what fill that blueprint after training, and they change with every release.
By freezing the blueprint and leaving the weights loadable, Frozen v2 keeps a future Gemini running on the same hardware. How much of the architecture will be hardcoded has reportedly not been settled by the engineers still working on the design.
Frozen v2 Compared With Google’s TPU Line
| Attribute | TPU 8t / TPU 8i | Frozen v2 (reported) |
|---|---|---|
| Purpose | 8t for training, 8i for inference, across many models | Gemini-family inference only |
| Flexibility | Runs any model loaded onto it | Tied to Gemini’s structural blueprint |
| Efficiency claim | TPU 8i: 80% better inference performance per dollar than predecessor (Google’s own figure) | 6–10x tokens per watt versus latest TPUs (internal engineering projection) |
| Availability | Sold and leased to external customers including Anthropic and Meta | Almost certainly internal only |
| Production scale | Volume product, part of Google Cloud backlog | Well below TPU volumes; treated internally as exploratory |
| Status | Announced at Google Cloud Next, April 2026 | Reported target of 2028; unconfirmed by Google |
The Compute Shortage Behind the Design
Google is not chasing this design for elegance. Running short of compute has been costing the company revenue it has already booked.
On the first-quarter 2026 earnings call, chief executive Sundar Pichai told analysts the company was “compute constrained in the near term” and that cloud revenue would have been higher had it been able to meet demand. Google Cloud revenue rose 63% year over year to about $20 billion that quarter, while the backlog of signed but undelivered contracts roughly doubled to approximately $462 billion.
The squeeze showed up externally and internally. Google told Meta it could not supply the full Gemini capacity Meta wanted to buy. In June 2026 it agreed to pay SpaceX roughly $920 million a month for access to around 110,000 Nvidia GPUs as bridge capacity for Gemini Enterprise — a company spending well over $180 billion on its own infrastructure was still renting close to a billion dollars of someone else’s compute every month. That arrangement sits alongside the broader consolidation of compute and energy assets under SpaceX. Inside DeepMind, researchers reportedly queued for the same TPUs Google was selling to Anthropic under a deal covering up to one million chips.
Talent moved with the compute. In a single week in June 2026, Noam Shazeer — co-lead on the Gemini models and co-author of the 2017 transformer paper — announced a move to OpenAI, while Nobel laureate John Jumper announced a move to Anthropic.
A chip that serves Gemini at several times the efficiency of a standard TPU multiplies the effective output of every facility Google already runs, without new grid connections, new floor space, or new fab slots. Given that electricity has become the binding constraint on AI expansion, that is the most valuable form of extra capacity available.
What Changes for Gemini Users and for Google’s Margins
Nothing about Frozen v2 alters what Gemini can do. It alters what Gemini costs to run — and at Google’s volume, cost is strategy.
Alphabet’s second-quarter 2026 results, reported on 22 July, put the scale in view. Google Cloud revenue reached $24.77 billion, up 82% year over year. The Gemini app crossed 950 million active users, processing 22 billion API tokens per minute, up from 16 billion a quarter earlier. Pichai said nearly 90% of Fortune 100 companies were using Gemini Enterprise, described Gemini 4 as “a very ambitious effort,” and told analysts to expect new model releases at close to a monthly cadence.
Serving that much traffic on general-purpose silicon is expensive. Inference already accounts for roughly two-thirds of AI compute spending industry-wide, and the share grows as deployed products scale. Cutting the energy per token by even a fraction of the projected amount hands Google room to price Gemini below rivals still renting general-purpose hardware, or to keep prices flat and bank the margin.
The spending picture makes the urgency plain. Alphabet raised full-year 2026 capital expenditure guidance to between $195 billion and $205 billion, up from $180 billion to $190 billion, after spending $44.9 billion in the second quarter alone — double the prior year. Finance chief Anat Ashkenazi told analysts the company remains “still in a supply-constrained environment.”
The Bet Underneath the Engineering
Committing a model’s architecture to silicon is a wager that the architecture has stopped moving.
The chip industry has resisted exactly this for years. General-purpose accelerators dominate because nobody could promise what model designs would look like three years out. Nvidia sells partly on the assurance that its GPUs will run whatever arrives next. Amazon’s Trainium, Microsoft’s Maia, and Meta’s MTIA are all more specialised than GPUs while staying neutral enough to host multiple model families.
Google is making the opposite claim: that the transformer structure underneath Gemini has settled enough that locking it into hardware costs less than the efficiency it buys. The evidence is not weak. Multi-head self-attention, feed-forward layers, residual connections, and normalisation placement have proven durable since 2017 across enormous variation in scale and training method.
The parallel that gets drawn most often is Apple’s silicon strategy, where chips are co-designed with the software that runs on them. Frozen v2 takes that further, designing not for a category of workloads but for one model family’s shape. It also inherits the same dependency chain that every custom accelerator programme faces on materials and fabrication.
What Could Go Wrong
The clearest failure mode is architectural drift. If Gemini were rebuilt around a state-space model, a hybrid design, or some successor paradigm, the hardwired portions would no longer match the model and the chip would lose part or all of its advantage regardless of how cleanly the weights update.
Commercial limits apply too. Because it cannot run other customers’ models, Frozen v2 is unlikely to appear as a Google Cloud product. It stays an internal efficiency tool, which caps the return to whatever Google saves on its own serving costs. Production volumes are expected to fall well short of TPU levels, and the effort is treated internally as a trial run.
Then there is the state of the evidence. The 2028 date is a reported goal. The efficiency numbers are engineering projections rather than measured results. Google has not publicly detailed the architecture, the process node, or the packaging. The report also landed during a difficult stretch for Google’s model roadmap, with Bloomberg reporting that Gemini 3.5 Pro had slipped behind schedule after missing internal targets on coding.
Timeline of the Frozen v2 Story
| Date | Event |
|---|---|
| April 2026 | Google announces TPU 8t and TPU 8i at Google Cloud Next; Pichai says the company is compute constrained. |
| Q1 2026 | Google Cloud revenue reaches ~$20 billion, up 63%; backlog roughly doubles to ~$462 billion. |
| June 2026 | Google agrees to pay SpaceX about $920 million per month for roughly 110,000 Nvidia GPUs; senior AI researchers depart for OpenAI and Anthropic. |
| 20 July 2026 | The Information reports Frozen v2; Alphabet stock rises as much as 3.7% intraday, closing up 1.51%. |
| 22 July 2026 | Alphabet Q2 results: Cloud up 82% to $24.77 billion; capex guidance raised to $195–205 billion. |
| 2028 (target) | Reported deployment window for Frozen v2. |
What to Watch Next
Three signals will tell you whether this becomes real hardware or stays a research exercise. The first is disclosure: Google has said nothing official, and any confirmation of architecture, node, or timeline would move the project from rumour to roadmap. The second is Gemini’s own design trajectory — a Gemini 4 or Gemini 5 that keeps the transformer skeleton intact strengthens the case, while an architectural departure weakens it. The third is imitation. If Frozen v2’s numbers hold up in production, every hyperscaler running a flagship model faces the same question about whether its own architecture has stabilised enough to commit to silicon.
For now, the honest summary is that Frozen v2 powers Gemini by making each answer cheaper rather than better. That is a smaller claim than the headlines suggest and a larger one than it sounds, because at 22 billion tokens a minute, the cost of an answer is the business. Whether the projected gains survive contact with fabrication, and whether the wider AI infrastructure stack moves the same direction, remains open for at least two more years.
If you are interested in this topic, we suggest you check our articles:
- What Is a Tensor Processing Unit (TPU)? A 2026 Guide
- How Much Electrical Power Does AI Require?
- AI Infrastructure: Essential Components in Modern ML Systems
- Essential Materials for AI Chip Production and Manufacturing
- Google Gemini: How Has It Been Received by Users So Far?
Sources: TechCrunch, CNBC, The Decoder, Tech Times, Quartz, CNBC Q2 earnings coverage, Investing.com earnings transcript
Written by Alius Noreika

