It could be TurboQuant or something like it. The biggest detraction to local LLM models is being able to close the gulf between obscenely-expensive 512GB NPUs, to house the 230GB uncompressed models (+ context), and more common 24GB GPUs. Quantized 15-18GB models are already working pretty well, but context size is still a bit of a problem.
Of course, the whole industry need to ramp up memory production and wrestle duopolies from the few that can make the raw silicon. It was pretty fucking pathetic that parts of the PC industry decided to leave these silicon processing weaknesses in various places. Large corps could have easily jumped into the industry and made bank in the long-term, but that would require not funneling into short-term quarterly profit bullshit.
It could be TurboQuant or something like it. The biggest detraction to local LLM models is being able to close the gulf between obscenely-expensive 512GB NPUs, to house the 230GB uncompressed models (+ context), and more common 24GB GPUs. Quantized 15-18GB models are already working pretty well, but context size is still a bit of a problem.
Of course, the whole industry need to ramp up memory production and wrestle duopolies from the few that can make the raw silicon. It was pretty fucking pathetic that parts of the PC industry decided to leave these silicon processing weaknesses in various places. Large corps could have easily jumped into the industry and made bank in the long-term, but that would require not funneling into short-term quarterly profit bullshit.