• [object Object]@lemmy.ca
    link
    fedilink
    English
    arrow-up
    19
    ·
    7 hours ago

    They’re oversold though, especially prompt caching and the parameter count war

    The US model is that they think more training compute and parameters will result in the winning model, while the Chinese are focusing on RL and parameters efficiency due to compute limits.

      • Dave.@aussie.zone
        link
        fedilink
        English
        arrow-up
        7
        ·
        edit-2
        4 hours ago

        Always ready to try brute force first. And then some other, less palatable options if that doesn’t work, like slightly less brute force.