ai-tools

Google is working on a new AI chip designed to make Gemini more efficient

vybecodingBy vybecoding.ai Editorial
July 20, 20265 min readOfficial
Google is working on a new AI chip designed to make Gemini more efficient
Google disclosed through its engineers, cited by The Information on July 20, 2026, that it is developing a custom server chip codenamed "Frozen v2" — a design intended to make Gemini models dramatically cheaper to run.

Google disclosed through its engineers, cited by The Information on July 20, 2026, that it is developing a custom server chip codenamed "Frozen v2" — a design intended to make Gemini models dramatically cheaper to run. Multiple reports indicate the chip could deliver six to ten times more AI output per watt than Google's current hardware, with a deployment target as early as 2028.

What's Converging

The announcement lands at a moment when every major AI lab is reckoning with the same uncomfortable truth: running frontier models at scale is expensive enough to threaten the economics of the entire business. Google's AI capital expenditure is expected to reach somewhere between $180 and $190 billion — a number that requires efficiency gains, not just revenue growth, to justify to shareholders. That pressure is now translating directly into silicon.

The race to build custom inference hardware is no longer a Google-specific story. OpenAI moved first among the hyperscalers-turned-model-labs, reportedly developing its own inference chip internally codenamed "Jalapeño" as recently as June 2026. Anthropic, which has no hardware division, is reportedly in talks with Samsung to explore a similar path. Our read is that this convergence matters more than any single announcement: when all three frontier labs independently decide the same vendor relationship — in this case, reliance on Nvidia — is a structural risk worth addressing with billions in engineering time, the underlying problem is real, not speculative.

Nvidia's GPUs remain the default infrastructure for training and serving large language models, and that dependency has consequences. GPU supply constraints have created internal friction at Google that, according to reporting from The Information and corroborated by Reuters, became acute enough to force Google Cloud to turn away outside customers. A company that cannot serve paying customers due to compute constraints has a different kind of problem than a company navigating ordinary capacity planning — it is leaving revenue on the table while simultaneously subsidizing its own model usage at high unit cost.

The Specific Development

The chip Google is building, internally called "Frozen v2," takes a fundamentally different approach from the company's existing Tensor Processing Units. Where TPUs are general-purpose AI accelerators that run whatever model is loaded onto them, Frozen v2 would permanently embed portions of Gemini's architecture directly into the silicon. The "frozen" naming is deliberate: locking model parameters into hardware eliminates the overhead of loading and computing those weights dynamically, which is where significant energy and latency costs accumulate during inference.

According to The Information's reporting — cited across TechCrunch, Reuters via Yahoo Finance, and The Economic Times — Google engineers project the chip could serve six to ten times more tokens per unit of power compared to the company's newest TPUs. That is not an incremental improvement. A 6x floor on efficiency gains, if realized at production scale, would fundamentally change the unit economics of running Gemini in consumer products and developer APIs alike. CNBC reported Alphabet shares climbed following the news, though the final close of roughly 1.5% up was more modest than the early-session pop of around 3.3% tracked by Reuters. The restrained close suggests markets welcomed the news but are pricing in the 2028 timeline appropriately — this is a research-and-design commitment, not a product launch.

Critically, multiple reports confirm that Frozen v2 is intended to run alongside TPUs, not replace them. The distinction matters for developers trying to understand what this means for the Gemini API: Google is not abandoning its existing hardware stack. It is adding a specialized inference layer for the specific workloads — high-volume, latency-sensitive, Gemini-specific serving — where the economics of general-purpose accelerators are hardest to justify. Engineers are, per The Information, still finalizing how much of the model will be hardwired, which means the efficiency ceiling is not yet fixed.

What's Likely Next

The immediate question is whether Frozen v2 actually ships in 2028. Custom silicon projects at this scale routinely slip, and Google has a history of ambitious hardware timelines that stretched — the third-generation TPU, for instance, spent longer in validation than originally projected. What distinguishes this project from past TPU generations, at least in the public description, is that it is not a general AI accelerator competing on FLOPS. It is a purpose-built inference engine for a single model family. That narrower scope could actually make it easier to finalize and manufacture, but it also makes it more brittle: any significant architectural shift in Gemini between now and 2028 potentially invalidates assumptions baked into the silicon.

Watch Google's next earnings call closely, which lands this week — Alphabet management will almost certainly face questions about the report, whether they confirm Frozen v2 by name or not. The more telling signal will be whether they frame efficiency gains as a 2027-2028 story or something they expect to start realizing sooner through software and existing TPU optimization. A parallel thread worth tracking: the Anthropic-Samsung talks. If Anthropic — which has less capital than Google and no existing chip infrastructure — commits to a custom hardware path, it signals that the Nvidia dependency problem is severe enough to justify the investment even without Google-scale resources. That outcome would be the strongest external validation that what Google is doing with Frozen v2 is a structural bet on the direction of AI infrastructure, not an internal engineering project with uncertain commercial payoff.

Sources

techcrunch.com Alphabet stock pops on report it's developing a more efficient AI chip Google plans new chip to run Gemini models more efficiently, the Information reports Google plans new chip to run Gemini models more efficiently: Report - The Economic Times

Based on

https://techcrunch.com/2026/07/20/google-is-working-on-a-new-ai-chip-designed-to-make-gemini-more-efficient/techcrunch.com

This article is an original, AI-assisted summary and analysis. Credit for the underlying reporting or footage belongs to the source above.

vybecoding

Written by the vybecoding.ai editorial team

Published on July 20, 2026

TOPICS

#ai#news
Google is working on a new AI chip designed to make Gemini more efficient