Google's 'Frozen v2' Chip Targets 6–10x Efficiency Gains for Gemini Inference by 2028

Alphabet is developing a new custom server chip internally called "Frozen v2" that could deliver six to ten times the inference efficiency of Google's current AI silicon, according to a report first published by The Information on July 20, 2026. Reuters independently corroborated the efficiency figure, citing the same tokens-per-unit-of-power metric (Reuters).
The chip is reportedly slated for release sometime in 2028. According to CNBC, Frozen v2 would embed parts of Gemini's architecture directly into the silicon, a design choice that shifts work from software to fixed-function hardware in pursuit of throughput and energy savings.
Google did not directly confirm or deny the report. A spokesperson told TechCrunch: "Our teams are constantly researching and experimenting with new innovations to deliver maximum performance and efficiency for our users and customers. While not every project moves into production, this rigorous exploration is central to our full stack approach." The phrasing is consistent with Google's standard handling of pre-release hardware leaks, though the explicit caveat that "not every project moves into production" is worth noting given the 2028 timeline.
Alphabet shares climbed approximately 3% on the morning of July 20, 2026, following the report (CNBC; Bloomberg).
Frozen v2 would be the latest entry in a custom-silicon lineage that stretches back nearly a decade. Google's Tensor Processing Units have powered Gemini and other foundation models for years, according to the company's official blog. The most recent generation, Ironwood, was described by Google as "the first Google TPU for the age of inference," designed to power thinking, inferential AI models at scale (Google Blog). A subsequent post referred to Google's eighth-generation TPUs as "two chips for the agentic era" (Google Blog). Ironwood delivers more than 4x better performance per chip for both training and inference workloads compared to the prior TPU generation (Google Blog).
A 6–10x efficiency jump would be a steeper generational improvement than Ironwood's already substantial 4x per-chip gains. The metric matters because inference cost, not training cost, is increasingly the dominant line item for large-model operators. As models grow and agentic workloads multiply call patterns, the number of tokens served per watt directly determines unit economics.
Embedding model architecture directly into silicon is the more technically significant detail. Custom ASICs have long incorporated domain-specific functional blocks, but hardcoding elements of a specific model family's architecture into fixed silicon reduces flexibility: the chip is optimized for Gemini's design assumptions and less useful for other workloads. That tradeoff only makes sense if Google expects Gemini's architectural lineage to remain stable through the chip's operational lifetime. A 2028 deployment would carry that silicon for several years beyond its introduction.
The broader context is a compute arms race in which every hyperscaler is pushing custom silicon to reduce dependence on merchant GPUs. Google's TPU program is the most mature of these efforts, and Frozen v2 suggests the company is moving further along the specialization axis, from general-purpose AI accelerators toward silicon co-designed with a specific model family. Nvidia's GPUs retain the advantage of running virtually any model; Google's bet is that vertical integration of model and silicon yields efficiency gains that general-purpose hardware cannot match.
There is risk in that bet. If model architectures shift substantially before 2028, hardware hardwired to current assumptions could age poorly. Google's own spokesperson acknowledged that not every research project reaches production. A two-year horizon between leak and ship date also leaves room for scope changes, delays, or cancellation.
What the market reacted to on July 20 was a directional signal: Google is planning to deepen its vertical integration of AI silicon well beyond what current TPU generations represent. Whether Frozen v2 ships on schedule and delivers the reported efficiency gains will depend on factors the leaked details do not capture: node availability, packaging decisions, software stack maturity, and whether Gemini's architecture evolves in ways the silicon can still serve. For now, the report confirms that Google is investing in the next major step of a strategy it has pursued for years.


