No. of Recommendations: 21
> "Google... using their cheaper - but less efficient - chips"
> "Nvidia is pursing an alternate strategy using their more efficient chips"
That is totally back to front.
Google's TPU chips are far more power/cost efficient than NVIDIA GPUs per unit work, when it comes to 'modern AI' / LLM inference or training.
A factor of 2x, typically, in the equivalent generations.
They're also a lot easier to cool since they run at much lower TDPs per unit.
It's a big part of the reason why Google are doing it themselves, they can optimise for a task they want to do a lot of, instead of buying general purpose chips.
For example they can have one chip totally designed to do training; another designed to do inference. And nothing else.
Nvidia's chips are more flexible / general purpose. If someone comes up with 'a better way of doing modern AI' (has happened every few months in recent years), it's easy to change.
Or if you want to be able to resell spare capacity for any kind of use, NVIDIA chips are more suitable.
Google's chips are more hard-wired for a particular approach. So you might have to make a new TPU chip physically if there's a big shift in how AI is done e.g. datatype/architecture or whatever.
The most surprising thing for me, is that despite NVIDIA buying Mellanox a few years back, Google is well in the lead when it comes to hooking up lots of chips to each other at scale.
NVLink you're talking 72 GPUs interconnected at the high end.
Google Tensor chips you can go up to nearly 10000.
Source: well, me. You can look it up anywhere you like to double check*... but my professional background is this field. I have no position in either company or in any AI company.
TRS
* but e.g.
telnyx.com - Tpu vs GPU