Bytedance is training an AI model with up to ten trillion parameters, according to the Financial Times. That's three times the size of Moonshot's Kimi K3, currently the largest Chinese model.
Having run models locally, RAM use seems to be almost directly proportional to number of parameters. 8 Billion parameters requires approx 8GB of VRAM at 1/4 precision.
Therefore, if this pattern holds you somehow need 10 Terabytes of VRAM at 4K and 40 Terabytes at full precision.
I think I saw some estimates that Claude’s Opus models may be and Opus model equivalents may be at around 100B parameters (100-400GB VRAM).
for inference you’re only counting active parameters towards VRAM, and some labs / models don’t train at 32b precision, or even use the same precision for different parts of the network
Having run models locally, RAM use seems to be almost directly proportional to number of parameters. 8 Billion parameters requires approx 8GB of VRAM at 1/4 precision.
Therefore, if this pattern holds you somehow need 10 Terabytes of VRAM at 4K and 40 Terabytes at full precision.
I think I saw some estimates that Claude’s Opus models may be and Opus model equivalents may be at around 100B parameters (100-400GB VRAM).
TLDR its clear why RAM is so expensive.
for inference you’re only counting active parameters towards VRAM, and some labs / models don’t train at 32b precision, or even use the same precision for different parts of the network