AIToday

Google designs Gemini-specific chip for 6–10× efficiency gain, deployment from 2028

THE DECODER14h ago
Google designs Gemini-specific chip for 6–10× efficiency gain, deployment from 2028

Key takeaway

Google is developing a specialized server chip called Frozen v2 that bakes Gemini's AI model architecture directly into hardware, with the aim of achieving 6 to 10 times greater efficiency in serving AI responses compared to its current TPU chips. Scheduled for deployment starting in 2028, the chip is intended to reduce Google's internal compute costs and strengthen its competitive position against OpenAI and Anthropic—though it is unlikely to become a commercial product since it is tied specifically to Gemini's architecture.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Google is building an internal server chip called Frozen v2 that embeds Gemini's model architecture directly into silicon. The chip could be 6 to 10 times more efficient at serving AI responses than Google's current TPU chips, and Google plans to deploy it starting in 2028.

  • Why it matters

    In AI, inference cost optimization increasingly determines profit margins. Frozen v2 is designed to ease Google's internal compute strain and could let Google run powerful models at lower cost—a potential competitive advantage against OpenAI and Anthropic. Unlike Google's TPU line, which it leases to Meta and external cloud customers, Frozen v2 is built specifically for Gemini and unlikely to become a commercial product.

  • What to watch

    Google has not yet decided how much of Gemini's architecture will be hardcoded into the chip. The design—which embeds the model architecture rather than weights—allows new weights to be loaded, making it more flexible than an earlier approach that would have locked the chip to a single Gemini version.

In Depth

Google is developing a specialized server chip internally named Frozen v2 that embeds Gemini's AI model architecture directly into silicon. According to sources cited by The Information, the chip could be 6 to 10 times more efficient at serving AI responses than Google's current TPU chips. Google plans to deploy Frozen v2 starting in 2028 with a smaller production volume than its TPU line, viewing it as a test run for specialized chips rather than a high-volume product.

The design philosophy behind Frozen v2 differs fundamentally from Google's general-purpose TPUs. While TPUs are built to work with many models and are leased to Meta, offered to external cloud customers, and positioned through Google's TPU@Premises program as an alternative to Nvidia, Frozen v2 is tailored exclusively to Gemini. This specialization is achieved by baking portions of Gemini's model structure directly into the hardware. The name Frozen v2 follows AI terminology: in machine learning, "freezing" parameters means locking their values so they stop changing during training. With Frozen v2, part of the model gets permanently frozen into the chip itself, cutting down on compute steps and speeding up responses.

The chip's design history reveals why this approach was chosen. The original concept reportedly came from Jeff Dean, Google Deepmind's chief scientist, who proposed embedding model weights—the specific numerical settings that determine how an AI model responds to queries—directly into silicon. Google scrapped that approach because a chip containing fixed weights would only work with a single version of Gemini and would become outdated too quickly. Frozen v2 takes a more flexible path by embedding the model architecture, meaning the underlying blueprint rather than the tuned parameters. New weights can still be loaded onto the chip, preserving adaptability. However, the exact amount of architecture that will be hardcoded has not yet been decided.

Because Frozen v2 only works as long as Google maintains the same model architecture, it is unlikely to ever become a commercial product for outside customers. Instead, the chip is designed to address Google's internal compute crunch. In the competitive AI landscape, however, inference-cost optimization increasingly determines profit margins. If Frozen v2 delivers on its promise of 6 to 10 times greater efficiency, it could give Google a significant edge—allowing the company to run powerful models at lower costs and potentially taking market share from OpenAI and Anthropic. This internal efficiency gain, rather than external revenue, appears to be the strategic goal of the Frozen v2 project.

Context & Analysis

Google's Frozen v2 represents a strategic shift toward application-specific chip design aimed at reducing the company's own inference costs. The project originated from an idea by Jeff Dean, Google Deepmind's chief scientist, who initially proposed embedding model weights directly into silicon. Google abandoned that approach because it would have locked the chip to a single Gemini version and rendered it obsolete quickly. Frozen v2 addresses this by embedding only the model architecture—the underlying blueprint—rather than the tuned parameters, allowing new weights to be loaded and preserving some flexibility.

The distinction between a general-purpose TPU and Frozen v2 reflects a broader tension in Google's chip strategy. TPUs serve multiple models and customers, including Meta and external cloud clients, positioning them as a competitive alternative to Nvidia infrastructure. Frozen v2, by contrast, is optimized solely for Google's internal use with Gemini, making it unsuitable for commercial sale. This design choice signals that Google views inference-cost reduction as a critical lever for profitability in AI: as competition with OpenAI and Anthropic intensifies, the ability to run powerful models efficiently directly affects margins. If Frozen v2 achieves its promised 6 to 10 times efficiency gain, it could give Google a cost advantage that translates into lower prices or better margins—though the chip's narrower scope means it will not reshape Google's external TPU business or its relationship with customers like Meta.

FAQ

When will Frozen v2 be available?
Google plans to deploy Frozen v2 starting in 2028. It will have a smaller production volume than Google's TPU line.
Will Google sell Frozen v2 to customers like it does with TPUs?
No. Because the chip only works as long as Google keeps the same model architecture, it probably won't become a product for outside customers. Frozen v2 is meant to ease Google's internal AI compute capacity.
How is Frozen v2 different from Google's existing TPU chips?
Unlike TPUs, which work with many models, Frozen v2 embeds parts of Gemini's model architecture directly into the hardware. This reduces compute steps and speeds up responses, making it 6 to 10 times more efficient at serving AI responses than current TPU chips.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime

1 minute a day. The AI essentials.

200+ sources · Email / LINE / Slack

Get it free →