google

Gemini etched into silicon: the "Frozen v2" chip claim

Promtime

google

Alphabet is developing a custom inference chip called Frozen v2 that would be 6 to 10 times more efficient than its current AI chips and could go into use as early as 2028, according to Stansberry Research's Weekend Edition, which cites a report from The Information.

At a glance

  • Architectural elements of the Gemini model would be etched permanently into the silicon, a design meant to cut how far data travels inside the chip and how many calculations each response requires.
  • Alphabet's expected capital expenditures this year approach $205 billion, while the newsletter puts cash from operations at $165 billion last year and close to $200 billion expected this year.
  • Every prompt requires a fresh round of calculations that consume computing power, electricity, memory and networking capacity, so cheaper inference silicon would reduce the recurring cost of each AI interaction.

Training is a bounded expense that ends when a model is finished, while inference recurs with every prompt, which is what makes silicon aimed at that step consequential. Alphabet appears to be treating cost per answer, rather than model capability, as the axis of the next competitive round. A sustained efficiency advantage there would likely surface as pricing room and margin protection rather than as benchmark headlines.

Frozen v2 etches part of the Gemini model into the physical silicon

Inference is the industry term for generating a response after a user prompts a model. Training happens for a limited period, while inference runs continuously across billions of daily requests. The Weekend Edition says Frozen v2 is designed specifically for that step.

The name refers to etching part of the model permanently into the physical silicon, which the report says minimizes data movement and reduces the volume of calculations needed per response. Chips built that way should draw less electricity, respond faster and handle more simultaneous requests, according to the newsletter.

AI engineers already talk in terms of inference efficiency and cost per token, a token being the smallest unit of text a model processes, around four characters, and the Weekend Edition argues competition will shift toward the cost of producing one more answer rather than PhD-level reasoning.

Alphabet expects capital expenditures approaching $205 billion this year

The newsletter puts Alphabet's expected capital expenditures this year at close to $205 billion. It cites cash generation of $165 billion from operations last year, with close to $200 billion expected this year, and says few companies can afford investment at that level. The Weekend Edition describes the capital spending figure as one that would have seemed almost unimaginable a few years ago.

Advertising accounted for more than 75% of Alphabet's sales last year, most of it generated by the Google search engine. The company now runs AI Overviews and Gemini-powered search features that answer questions, summarize information and compare products instead of returning links alone. According to the Weekend Edition, better recognition of search intent lets Google show more relevant ads.

Google Cloud sales grew 82% last quarter

Google Cloud generated nearly $60 billion in revenue last year, around 15% of Alphabet's sales, and grew 36% over the year, close to double the 20% growth the newsletter cites for Amazon's and Microsoft's cloud businesses. Growth reached 82% last quarter.

The Gemini app has 950 million monthly active users, and Alphabet's models process 22 billion tokens per minute. Anthropic, a Google Cloud customer, has committed to $200 billion in Google Cloud spending over the next five years, according to the newsletter.

The Weekend Edition frames AI as one of the most expensive computing projects ever undertaken, saying OpenAI, Anthropic and xAI have lost tens of billions of dollars building competing models. It adds that AI model startups cannot afford their own large data centers and rent Alphabet's infrastructure instead.

What the 2028 date leaves open

The earliest deployment named is 2028, and the newsletter gives no manufacturing partner, production volume, per-unit cost, or indication of how much of Gemini's architecture the design would lock into silicon. The 6 to 10 times efficiency figure comes from the report rather than from Alphabet, and no benchmark or workload is attached to it.

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.