

lol might have to smuggle it directly from China 🤣


Sure, a benchmark doesn’t capture all the subtleties and different use cases, but it does give a general idea of the capabilities of a model. Obviously, you have to run the model and see if it does what you need. But the chart isn’t really about the nuance, it’s showing how drastically the efficiency of the models has improved in just a year. The fact that we can even reasonably compare a model you can run on a desktop to one that needed a data center just a year ago is phenomenal.


ah gotcha, and I’ve made the list myself apparently


I’m fairly optimistic that people will figure out how to optimize the models a lot further going forward. One obvious path is to try and separate the reasoning network from the trivia that gets baked into the model, and some work is being done in this area. If you could have a context free reasoning engine and then feed the facts it needs to know on the fly based on the context you’re running it in, then you could likely have a much smaller model that’s very capable.


Not sure what Moore’s law has to do with anything here to be honest. The models you can run locally on a consumer desktop can do real work, and their resource usage is no different from any other software like games that you’d run.


The article doesn’t really talk about this directly, but when you read it critically it becomes clear that capitalist corporate structures are ill equipped to deal with the negative effects of LLMs.


You need a GPU with around 16gb vram at a minimum to run qunatized version.


that’s the other huge advantage of open models you can run locally


The difference is that you can run Qwen completely local though.


yeah, it’s not a completely insane amount of data, and a db like postgres can do fast text search on that too with fuzzy matching


Ok, but that’s a completely nonsensical statement. If you ever used Qwen in an agentic loop, you’d know that it delivers working code, and it takes about same resources as playing a modern game, and I don’t see anybody whinging that game are too inefficient for what they deliver.


I mean baking knowledge into a model isn’t really all that useful to begin with. Just download wikipedia locally and have it access it through tool use, it’s way more efficient and more accurate. And yeah, I find Q6 tends to be the sweet spot where it’s close enough to full 16 bit in performance, but doesn’t chew up too much memory.


some turbolib made a Lemmy client that had a hardcoded blacklist of users and instances they considered to be communists


Ok that’s fair, what Iran is doing is very similar to Russia’s approach of launching combination of drones and missiles to overwhelm Ukrainian defences.
Seems like the worker movement is a lot stronger in Bolivia, and they actually ran US puppet out of the country already.


I don’t really see the misleading part in the headline myself. While China and Russia undoubtedly help Iran, it seems that Iranians are perfectly capable to inflict massive damage on the empire using their own technical capabilities. While Ukraine is a western proxy entirely dependent on its western masters, Iran has its own deep technological and industrial base. It’s worth remembering that it was Iran that originally did a technology transfer of their drone technology to Russia rather than the other way around.
Honestly, I think the most reasonable approach is just to see what other people’s experience is like and which models are well regarded, then try them out and see which one is the best fit for what you’re doing. You might not even need the top performing one necessarily, and speed or lower resource usage might be a bigger factor.