Google's Tensor Processing Unit production could reach 9 million units by 2027, roughly triple current volumes, according to projections cited in the market. The expansion would put the search giant on a collision course with NVIDIA in the AI hardware market.
The figure is a projection, not a confirmed order book. But even a partial ramp of that size would change the landscape of AI chips, an arena NVIDIA has dominated since the AI boom began. Google builds TPUs primarily for its own cloud and internal workloads, but it also offers them to customers through its cloud platform.
Why the ramp matters
TPUs are custom ASICs, purpose-built for tensor operations, which underpin machine learning models. Unlike NVIDIA's general-purpose GPUs, they're narrower in scope, but they run certain workloads with far less power and cost per inference. That makes them attractive for large-scale AI serving, the part of AI that runs after a model is trained.
Tripling volume to 9 million units in roughly two years isn't a small adjustment. It means fabricating more chips, securing more packaging capacity, and lining up memory suppliers. It also means Google believes there is enough demand from its own products and external cloud customers to justify the spend.
What this means for NVIDIA
NVIDIA still holds the top spot in AI accelerators, with a datacenter GPU business that brings in tens of billions per quarter. But the company's dominance has already attracted a wave of challengers — from startups to cloud giants building custom silicon. Google's scale makes it one of the most serious.
Google isn't just a chip designer, it's also the operator of the infrastructure that runs the chips. That dual role gives it a price advantage it can pass to customers, and it can design silicon around its own models and services. NVIDIA, by contrast, sells to everyone but designs for a broader set of workloads.
There's a real question about whether 9 million units will actually ship. Google doesn't buy TPUs from an external foundry, it designs them and works with manufacturers. The number could reflect internal forecasts, but it could also be a ceiling, not a target.
What's driving the push
The push comes down to the cost of running AI at scale. As models grow and get used more, the biggest expense shifts from training to inference. Google has been vocal about the total cost of ownership advantages of TPUs for serving — lower power draw, better utilization, no need to buy a high-end GPU for simple requests.
There's also a strategic reason: Google wants to control the whole stack, from model to chip to datacenter. It has already used TPUs to run its own Gemini models, and making more of them lets it scale AI products without renting capacity from rivals.
For NVIDIA, this isn't an overnight threat. GPUs are still the default for training, and NVIDIA's software ecosystem is a major lock-in. But if TPU volumes keep climbing, and if Google keeps winning large cloud deals on price, NVIDIA's lead in AI could start to erode.
All of it depends on whether Google can actually deliver. Triple the volume in two years is a tall order, and any hiccup — a supply shortage, a delay in the next TPU generation — would slow the momentum. For now, the numbers suggest Google is serious. The rest is execution.



