That’s also very possible. The US grid has very little spare capacity, and building out more will be a decades long project. So, if their newer models are more power hungry, then they might not be economically viable even with all the investor money being thrown at them.
I expect more power efficient chips that are ai specific to come out in the next few years. Eventually you’ll be able to run good models on your phone. Not sure about ram requirements or anything like that if the model could be shrunk down somehow. There’s definitely huge gains in optimizing efficiency to be had. Right now is the equivalent of an old IBM mainframe trying to do a spreadsheet. We might even giggle at the thought of gigabytes of ram in the future with having multiple terabytes as standard on personal devices.
I expect we’ll start seeing stuff like Taalas where they print the model to the chip and other specialized chips like Xuantie C950 going forward. Neither of these requires DRAM, and Taalas is particularly clever since they just print the model right to an ASIC chip. So, the whole renting out LLMs business model isn’t going to last long I suspect.
So a repeat of the crypto crash for graphics cards when ASICs ate their lunch. Mind you, that’s only for inference (although a super fast QWEN 3.8 would meet a lot of peoples needs).
The argument for datacentres is for training the models, but then they’ll need to prove that they haven’t hit a diminishing returns wall, which will be hard if, as seems likely, they have. Also the Chinese have been doing it in a cave, with a box of scraps (figuratively), and gotten at least 90+% as good results.
Seems like the recent advances have been in the frameworks, which don’t need no stinking (literally if fossil fueled) datacentres.
Musk: Yes, we should all slow down. With no external verification and upon the agreement of this handshake, we should all stop developing so fast. We, especially, will slow down. You can trust us.
Actually, what does that look like? I assumed from the expansions the bottleneck wasn’t algorithmic, i.e. each data center that they bring online was designed from the ground up to be at 100% all the time. Is that not the case? Can you even (for lack of a better metaphor) underclock your datacenter? Does the cooling work that way?
For that matter, are they doing that trick where the datacenter is owned/operated by Independent DC Company X, and has exclusive lease agreements for compute?
Maybe they are running out of electricity / data centres / some other requirement?
That’s also very possible. The US grid has very little spare capacity, and building out more will be a decades long project. So, if their newer models are more power hungry, then they might not be economically viable even with all the investor money being thrown at them.
I expect more power efficient chips that are ai specific to come out in the next few years. Eventually you’ll be able to run good models on your phone. Not sure about ram requirements or anything like that if the model could be shrunk down somehow. There’s definitely huge gains in optimizing efficiency to be had. Right now is the equivalent of an old IBM mainframe trying to do a spreadsheet. We might even giggle at the thought of gigabytes of ram in the future with having multiple terabytes as standard on personal devices.
I expect we’ll start seeing stuff like Taalas where they print the model to the chip and other specialized chips like Xuantie C950 going forward. Neither of these requires DRAM, and Taalas is particularly clever since they just print the model right to an ASIC chip. So, the whole renting out LLMs business model isn’t going to last long I suspect.
So a repeat of the crypto crash for graphics cards when ASICs ate their lunch. Mind you, that’s only for inference (although a super fast QWEN 3.8 would meet a lot of peoples needs).
The argument for datacentres is for training the models, but then they’ll need to prove that they haven’t hit a diminishing returns wall, which will be hard if, as seems likely, they have. Also the Chinese have been doing it in a cave, with a box of scraps (figuratively), and gotten at least 90+% as good results.
Seems like the recent advances have been in the frameworks, which don’t need no stinking (literally if fossil fueled) datacentres.
If so it amuses me that investing in maintaining and upgrading public infrastructure via taxes might have saved them the choke point
Musk: Yes, we should all slow down. With no external verification and upon the agreement of this handshake, we should all stop developing so fast. We, especially, will slow down. You can trust us.
Actually, what does that look like? I assumed from the expansions the bottleneck wasn’t algorithmic, i.e. each data center that they bring online was designed from the ground up to be at 100% all the time. Is that not the case? Can you even (for lack of a better metaphor) underclock your datacenter? Does the cooling work that way?
For that matter, are they doing that trick where the datacenter is owned/operated by Independent DC Company X, and has exclusive lease agreements for compute?