• MangoCats@feddit.it
    link
    fedilink
    English
    arrow-up
    9
    ·
    edit-2
    24 hours ago

    I see a business model where “we’re done, this one is (finally) good enough and now we’ll stop bleeding cash on the training and turn up the screws on the customers we’ve hooked on loss leader pricing.” Open weight models will never stop training for improvement, the costs for training seem to be inexorably falling, and any business model built on the idea that they can kick back and roll in the profits after their initial “hard work” is going to lose all their customers to better products.

    This isn’t some captive market like US automobile customers who have no choice but the limited selection of crap, crap and more crap that is put in front of them. At least not as long as the internet remains relatively open.

    • boonhet@sopuli.xyz
      link
      fedilink
      English
      arrow-up
      1
      ·
      edit-2
      8 hours ago

      It’s not necessarily that any of the open models are cheaper to train (for similar model size), it’s more that China has deeper pockets. And MoE inference is cheaper than dense models I believe

      • MangoCats@feddit.it
        link
        fedilink
        English
        arrow-up
        1
        ·
        4 hours ago

        If China manages to scale their hardware production, this could finally be “their day in the sun” where they clearly surpass the West in the way the West has outshone them for 100 years. MoE has its place, but a MoDE where each E is itself a dense model would be more powerful still, mostly you need the silicon gates and power to drive them.

        Three Gorges makes hydro-power, right? 22MW -> 100 TWh per year, just from that one structure, I bet that will run at least one AGI… https://www.fredgao.com/p/deepseeks-liang-wenfeng-breaks-his