• Eager Eagle@lemmy.world
    link
    fedilink
    English
    arrow-up
    3
    ·
    2 days ago

    for inference you’re only counting active parameters towards VRAM, and some labs / models don’t train at 32b precision, or even use the same precision for different parts of the network