I see a business model where “we’re done, this one is (finally) good enough and now we’ll stop bleeding cash on the training and turn up the screws on the customers we’ve hooked on loss leader pricing.” Open weight models will never stop training for improvement, the costs for training seem to be inexorably falling, and any business model built on the idea that they can kick back and roll in the profits after their initial “hard work” is going to lose all their customers to better products.
This isn’t some captive market like US automobile customers who have no choice but the limited selection of crap, crap and more crap that is put in front of them. At least not as long as the internet remains relatively open.
It’s not necessarily that any of the open models are cheaper to train (for similar model size), it’s more that China has deeper pockets. And MoE inference is cheaper than dense models I believe
If China manages to scale their hardware production, this could finally be “their day in the sun” where they clearly surpass the West in the way the West has outshone them for 100 years. MoE has its place, but a MoDE where each E is itself a dense model would be more powerful still, mostly you need the silicon gates and power to drive them.
I see a business model where “we’re done, this one is (finally) good enough and now we’ll stop bleeding cash on the training and turn up the screws on the customers we’ve hooked on loss leader pricing.” Open weight models will never stop training for improvement, the costs for training seem to be inexorably falling, and any business model built on the idea that they can kick back and roll in the profits after their initial “hard work” is going to lose all their customers to better products.
This isn’t some captive market like US automobile customers who have no choice but the limited selection of crap, crap and more crap that is put in front of them. At least not as long as the internet remains relatively open.
It’s not necessarily that any of the open models are cheaper to train (for similar model size), it’s more that China has deeper pockets. And MoE inference is cheaper than dense models I believe
If China manages to scale their hardware production, this could finally be “their day in the sun” where they clearly surpass the West in the way the West has outshone them for 100 years. MoE has its place, but a MoDE where each E is itself a dense model would be more powerful still, mostly you need the silicon gates and power to drive them.
Three Gorges makes hydro-power, right? 22MW -> 100 TWh per year, just from that one structure, I bet that will run at least one AGI… https://www.fredgao.com/p/deepseeks-liang-wenfeng-breaks-his