Houston News Buzz

collapse
Home / Daily News Analysis / Google is working on a new AI chip designed to make Gemini more efficient

Google is working on a new AI chip designed to make Gemini more efficient

Jul 24, 2026  Twila Rosenbaum 10 views
Google is working on a new AI chip designed to make Gemini more efficient

Alphabet, Google's parent company, is designing a new server chip to help its in-house Gemini models operate more efficiently. The chip, internally dubbed 'Frozen v2,' is slated to be released sometime in 2028, according to a report by The Information citing anonymous sources. The chip could be between six and ten times more efficient than Google's existing AI chips, measured by the number of tokens generated per unit of power. This efficiency gain is critical as AI companies face mounting pressure to reduce costs and improve performance.

In a response to TechCrunch, the company didn't directly confirm the report but didn't deny it either. 'Our teams are constantly researching and experimenting with new innovations to deliver maximum performance and efficiency for our users and customers,' Google told TechCrunch. 'While not every project moves into production, this rigorous exploration is central to our full stack approach. By co-designing our hardware and software from the ground up, we ensure our systems are integrated and highly optimized for real-world workloads.'

AI companies have increasingly sought to produce their own chips as a way to make their in-house models run more efficiently and to address global shortages in AI computing capacity. Such efficiency has become a key selling point for tech companies as concerns about AI spend have dampened the market euphoria that previously characterized the industry. At the same time, firms are engaged in an ongoing attempt to wean themselves off chipmaker Nvidia, which has historically dominated the AI chip market and whose dominance has left major AI makers dependent on its hardware.

The new Frozen v2 chip represents a significant leap for Google's custom silicon efforts. Google has been designing its own custom Tensor Processing Units (TPUs) since 2015, with the latest generation, the TPU v5p, being used for both training and inference. The Frozen v2 appears to be a separate line focused specifically on inference workloads for the Gemini family of models. By making the chip more efficient, Google can reduce the power consumption and cost of running these massive models, which is crucial as the company plans to spend between $180 billion and $190 billion on AI infrastructure this year alone.

Investors have previously worried about Alphabet's massive planned expenditures designed to help it build out its AI strategy. With so much money at stake, the company needs to prove that those investments will pay off. News of the more efficient Frozen v2 chip appears to have assuaged investors, giving Google a boost ahead of its earnings report later this week. Following publication of The Information's report, the company's stock climbed some 3% on Monday morning.

The drive for efficiency is not limited to Google. In June, OpenAI announced its first custom chip, an inference processor dubbed 'Jalapeño.' That chip is designed to optimize the performance of OpenAI's GPT models and reduce dependency on Nvidia hardware. Earlier this month, it was reported that Anthropic was discussing a new chipmaking partnership with Samsung. These moves highlight a broader trend in the AI industry: major players are increasingly investing in custom silicon to gain a competitive edge, lower costs, and ensure supply chain resilience.

Nvidia has long been the dominant force in AI chips, with its GPUs powering most AI workloads in data centers. However, the high cost and scarcity of Nvidia's hardware have pushed companies like Google, Amazon, and Microsoft to develop their own chips. Google's TPUs have been a key part of its cloud offering, but they have largely been used for Google's own models and services. With the Frozen v2 chip, Google aims to make Gemini more efficient not just for internal use but also for customers who use Google Cloud to run their AI workloads.

The Gemini family of models, which includes Gemini Ultra, Pro, and Nano, represents Google's latest push in generative AI. These models power everything from the Bard chatbot to search features and enterprise applications. Making Gemini more efficient could allow Google to offer lower prices for API access and cloud services, potentially undercutting competitors like OpenAI and Anthropic. Efficiency also means lower latency and better user experiences, which is critical in real-time applications like chat and search.

Industry analysts have noted that the reported 6-10x efficiency gain is ambitious but not unrealistic given the rapid pace of innovation in chip design. The improvement likely comes from architectural innovations tailored for transformer-based models, as well as process technology advancements. Google has deep expertise in both chip design and AI model development, which gives it an edge in co-designing hardware and software. The company's 'full stack approach' means it can optimize every layer of the stack, from the transistor level to the algorithm.

However, challenges remain. The chip is not expected to ship until 2028, which is several years away. By then, Nvidia will have released multiple new generations of its GPUs, potentially maintaining its lead. Moreover, Google has a history of canceling or delaying hardware projects. For instance, its custom server chips for general-purpose computing, codenamed 'Whitechapel,' were reportedly shelved. The Frozen v2 chip will need to survive internal budget reviews and continue to receive executive support to reach production.

Despite these uncertainties, the development of Frozen v2 signals Google's long-term commitment to AI hardware. The company is investing billions in its own chips and data centers, aiming to create a vertically integrated AI stack that rivals those of its biggest competitors. If successful, Frozen v2 could be a game-changer for Google's AI ambitions, enabling it to run Gemini models at a fraction of the cost of using Nvidia hardware. This could translate into lower prices for consumers and businesses, as well as more powerful AI capabilities.

In the broader context, the AI chip race is intensifying. Startups like Cerebras, Graphcore, and Groq are also developing specialized AI chips, while tech giants are building their own. The market for AI chips is expected to grow to over $300 billion by 2030, according to some estimates. Google's move to develop Frozen v2 is part of a strategy to capture a larger share of this market and reduce its dependency on external suppliers. The company has already demonstrated its ability to design competitive chips with its Tensor lineup for Pixel phones, and now it is applying similar expertise to data center chips.

The efficiency metric mentioned in the report—tokens generated per unit of power—is a crucial one for inference workloads. For large language models, the cost and energy consumption of generating tokens (pieces of text) are major concerns. A 10x improvement in token efficiency would mean that Google can serve the same number of queries with 90% less power, or serve 10 times more queries with the same power. This could dramatically lower the operational costs of running Gemini, making Google's AI services more profitable and scalable.

Google's stock rally following the news indicates that investors are optimistic about the potential of Frozen v2. However, the company still faces significant competition from Nvidia, as well as from custom chips being developed by Amazon (Trainium, Inferentia), Microsoft (co-developed with AMD), and others. The next few years will reveal which approach—general-purpose GPUs or custom AI chips—offers the best balance of performance, efficiency, and cost. For now, Google is betting that its deep integration of hardware and software will give it an edge in the race to dominate the AI chip market.

As the AI industry continues to mature, the importance of custom silicon cannot be overstated. The ability to design chips specifically for one's own models can unlock efficiency gains that are difficult to achieve with off-the-shelf hardware. Google's Frozen v2 chip, if it delivers on its promises, could be a key differentiator in the competitive landscape of AI. However, given the long timeline to 2028, the company must also continue to optimize its existing TPU lineup and collaborate with other chipmakers to meet immediate needs. The path to AI dominance is paved with silicon, and Google is digging deep to secure its supply.


Source:TechCrunch News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy