They’re oversold though, especially prompt caching and the parameter count war
The US model is that they think more training compute and parameters will result in the winning model, while the Chinese are focusing on RL and parameters efficiency due to compute limits.
They’re oversold though, especially prompt caching and the parameter count war
The US model is that they think more training compute and parameters will result in the winning model, while the Chinese are focusing on RL and parameters efficiency due to compute limits.
Reads like American vs European/Asian cars
Always ready to try brute force first. And then some other, less palatable options if that doesn’t work, like slightly less brute force.