You can run MiMo 2.5 in 128GB, as 3 bit quant. The model itself is less that 128GB as an IQ3_KT, and it’s KLD (measured quantization loss) is very reasonable.
I get about 9 tokens/sec on 128GB with a 7800 CPU and one RTX 3090. Intend to try dflash to speed it up this week. I can run up to 90k context, without too much kv quantization, depending on how I configure my PC.
You can run MiMo 2.5 in 128GB, as 3 bit quant. The model itself is less that 128GB as an IQ3_KT, and it’s KLD (measured quantization loss) is very reasonable.
I get about 9 tokens/sec on 128GB with a 7800 CPU and one RTX 3090. Intend to try dflash to speed it up this week. I can run up to 90k context, without too much kv quantization, depending on how I configure my PC.
And it’s a really good model, even quantized.