Skip to main content

Google wants to „engrave“ Gemini directly into the chip: The path to 10x AI efficiency

Ilustrační obrázek
Google is moving away from general-purpose processors towards extreme specialization. The "Frozen v2" project aims to link the architecture of the large language model Gemini with the physical structure of the chip, which could increase inference efficiency (calculations when generating responses) up to tenfold compared to current solutions.

In the race for dominance in artificial intelligence, the battle is shifting from the software level to the depths of hardware. While most companies focus on how to train larger models, Google has decided on a different move: it wants its Gemini models to be part of the very DNA of its chips. According to information from Tom's Hardware and Finance Biggo, Google is working on a project codenamed Frozen v2, which aims to change the rules of the game in AI operational efficiency.

From General Processors to "Flexible Hardware"

Until now, we have been used to AI models running on processors like GPUs (graphics processing units from Nvidia) or TPUs (Tensor Processing Units from Google). While these chips are powerful, they are "general-purpose." This means they must be prepared to process various types of tasks and models, which requires constant data transfer between memory and the processing unit. This process is energy-intensive and slow.

The Frozen v2 strategy is different. Instead of the chip merely "understanding" the model's instructions, Google wants to directly "etch" the Gemini architecture into the silicon. This minimizes the need for complex real-time decision-making processes. As a result, the chip doesn't figure out how to interpret the model but directly executes its mathematical operations with minimal latency.

A key element is the concept of "flexible hardwired". Unlike the original "hardwired everything" idea, which would make the chip obsolete after the first model update, Frozen v2 maintains flexibility. The architecture (the way data moves) will be fixed, but the model's parameters themselves (weights) will still be updateable. This allows the Gemini model to remain modern even when using specialized hardware.

Efficiency Comparison: Frozen v2 vs. Current Solutions

According to internal estimates by Google engineers, Frozen v2 could generate 6 to 10 times more tokens per unit of energy consumed than current generations of TPUs. For context, in the realm of large language models, "tokens per watt" is the most important indicator of cost-effectiveness.

Parameter Nvidia GPU (General) Google TPU (Specialized) Google Frozen v2 (Targeted)
Model Flexibility Extreme High Limited (Gemini only)
Inference Efficiency Standard High Extreme (6–10× higher)
Cost Optimization Medium High Maximum

Why is this happening now?

The reason is not just the pursuit of performance, but primarily a shortage of computing power. Google Cloud is currently facing enormous demand for AI capacities that exceeds the current supply. This has led to a situation where Google has had to turn away some customers. The specialized Frozen v2 chip is a direct response to this "bottleneck." If Google succeeds in reducing energy and hardware requirements for running Gemini, it will be able to serve many more users at the same cost.

We see this trend with competitors as well. Canadian startup Taalas is already working on a similar concept of hardware-encoding models. Giants like Microsoft and OpenAI are also investing in their own chips to reduce their dependence on Nvidia. Nvidia, on the other hand, is defending itself with strategic acquisitions (e.g., a license for Groq) and efforts to integrate into the entire ecosystem.

Practical Impact: What does this mean for you?

You might be thinking: "I'm not a data center, why should I care?" The answer is simple: price and availability.

  • For developers and companies in the Czech Republic: If you are building an application using the Gemini API (e.g., via Google AI Studio or Vertex AI), Google's lower operating costs potentially mean a lower price per token. This can be crucial for Czech startups that need to manage their budgets efficiently.
  • Availability in Czech: Gemini is already very well available in Czech. If Google can scale more cheaply thanks to new chips, we will see faster integration of advanced features (such as multimodal video analysis or a long context window) for our local needs.
  • Ecology and EU regulations: The European Union, through the EU AI Act, is placing increasing emphasis on the energy intensity of models. If Google can reduce energy consumption when generating responses by up to tenfold, its services will have a much easier path to compliance with European environmental standards.

Currently, Gemini is available through paid plans (e.g., Gemini Advanced for approximately 20 EUR/month) and free versions for developers. The transition to Frozen v2 could make these services even more accessible to the general public.

Conclusion

Google is not just trying to make its models smarter. It is trying to make them economically sustainable. If the Frozen v2 project succeeds (planned deployment around 2028), there could be a significant drop in the prices of AI services and, at the same time, a drastic increase in the speed we know today from common chatbots.

Will Frozen v2 replace existing TPU chips?

No. Google plans to develop these chips in parallel. TPUs will remain general-purpose tools for various types of AI tasks, while Frozen v2 will be a highly specialized "sprinter" designed exclusively for optimizing the Gemini architecture.

Can Frozen v2 be used for models like GPT or Claude?

No. The main advantage of the chip is precisely its specialization. The architecture is "etched" directly for Gemini. For other models, this chip would be inefficient and essentially unusable.

What is the price for using Gemini in the Czech Republic?

A free version is available for regular users. A Gemini Advanced subscription costs approximately 20 EUR per month (in the Czech Republic, the price is converted to CZK according to the current exchange rate). For businesses, prices are determined by token consumption within Google Cloud.

X

Don't miss out!

Subscribe for the latest news and updates.