No. of Recommendations: 9
Tokens per watt gains in efficiency are not all silicon based.
Many of the efficiency gains are from cooling, packaging, and software.
So older silicon can still increase tokens/watt.
Sure.
But that presumably doesn't help the *gap* between performance on older and newer hardware, presumably they would both benefit somewhat proportionately from better software. So older power-hungry units might still get put out to pasture.
Jim
* A bigger (but likely more remote) risk to the value of the infrastructure being built is that there might be a major change closer to the upper conceptual levels of the software. What if quite different computing setups will be needed for those new designs? I've long had the uncomfortable feeling that a lot of teams have jumped in with both feet on the current LLM designs and assumed that more is not only better, but is also best and will remain so: that simply "more"--more computing power, more parameters on existing models--will get them where they want to go better than any other technique that may come along in future. Given how fast things are changing, that seems like a risky assumption. Core memory was truly fantastic when it came out - (I've used it) - but imagine if someone had built a gigafactory to churn it out by the train load because it was obviously the best thing and more is better. What would those data centres stuff with carloads of core memory be worth right after the introduction of TTL? How fast can you say "other than temporary impairment"?