• i_am_not_a_robot@discuss.tchncs.de
    link
    fedilink
    English
    arrow-up
    5
    ·
    2 days ago

    Who is downloading these? I thought I would try out the new GLM before it becomes illegal, but these new models require 256GB RAM to run even the scaled down versions. I thought my PC I bought just before hardware prices went crazy was excessive but it only has 128GB. You’re looking at spending over 5000 dollars on a PC that can just barely fit one AI model and only until they make new ones that need 384GB.

    • brucethemoose@lemmy.world
      link
      fedilink
      arrow-up
      1
      ·
      edit-2
      22 hours ago

      You can run MiMo 2.5 in 128GB, as 3 bit quant. The model itself is less that 128GB as an IQ3_KT, and it’s KLD (measured quantization loss) is very reasonable.

      I get about 9 tokens/sec on 128GB with a 7800 CPU and one RTX 3090. Intend to try dflash to speed it up this week. I can run up to 90k context, without too much kv quantization, depending on how I configure my PC.

      And it’s a really good model, even quantized.

    • ☆ Yσɠƚԋσʂ ☆@lemmy.mlOP
      link
      fedilink
      arrow-up
      8
      ·
      2 days ago

      There are smaller models like Qwen 3.6 27b that can run with as low ast 16gb VRAM. The larger models are largely used either by companies running them on prem, or research institutions.