

Oh they definitely aren’t, there’s an interview with Alibaba Cloud founder where he discusses the direction in China. Basically, the goal is to find useful niches for this tech early on, then iterate and improve. They’re not chasing AGI or trying to make one model to rule them all. That said thoough, the capabilities of Chinese models in the same domains where American ones shine are very close as well. So, I do expect that Chinese models will catch up and start surpassing American ones on their own turf before long. I’m also expecting that the trend will shift towards running smaller and local models for most things because you just don’t need a giant model to do most tasks.


I don’t see how anything they do can possibly affect what Chinese labs are doing. And that’s the only alternative to American labs right now. So, who are they going to convince exactly?


I meant that simple making models bigger might not actually make them more capable. So even if you had unlimited hardware to play with, you might have to find a different approach.


I expect so as well, and my prediction is that we’ll have LLMs that are roughly as capable as the current frontier that can be run locally within a year or two. At that point, it’s just going to be good enough for vast majority of tasks most people need to do.


There’s no reason to think that the architecture itself can scale indefinitely. It might very well be that LLMs have some hard constraints on the scope of the problems they’re capable of solving.


That’s my view as well, LLMs are likely just one piece of a much bigger puzzle and we’re now hitting the limit of what you can do with them in practical terms.


Right, I’d argue that China proves you don’t need massive data centers for training. And yeah, I think something like Qwen 3.8 is more than enough for tasks most people do. There are a lot of tricks you can do as well with the harness, where there’s a lot of attention is shifting now. And it’s a lot cheaper and faster to develop better harnesses than train new models. I expect we’ll start seeing a shift towards neurosymbolic systems before long where the LLM acts as a stochastic component within a symbolic logic engine.


Again, the elephant in the room is that China is still pursuing this tech and they’re not signing up to any moratoriums. So, if they really believed this tech was dangerous and powerful, they’d be racing to develop it further before China does.


Right, they might just focus on big business, or even angle to become a vendor of record for the government. So, individual users might not really be of interest anymore.


There is absolutely no reason to expect that you can scale LLMs indefinitely.


They’d keep releasing them if there was money in it.


The whole bubble might be about to pop.


I expect we’ll start seeing stuff like Taalas where they print the model to the chip and other specialized chips like Xuantie C950 going forward. Neither of these requires DRAM, and Taalas is particularly clever since they just print the model right to an ASIC chip. So, the whole renting out LLMs business model isn’t going to last long I suspect.


yup, that’s a totally valid option if the costs for their bigger models are going through the roof


That’s also very possible. The US grid has very little spare capacity, and building out more will be a decades long project. So, if their newer models are more power hungry, then they might not be economically viable even with all the investor money being thrown at them.


In the end, the nazis won ww2 thanks to the US.


That’s why it’s essential to seize the means of vote production.
They have no leverage over Chinese labs, and China has every incentive to continue developing this tech. The only real explanation I see here is that they’re starting to get into diminishing returns territory, investors are getting edgy, and the costs of running this stuff are exploding.