In my regular research (behind a paywall), I have been saying for a while that I think the future of AI is not large language models (LLM), but small language models (SLM) run on local desktop computers or even mobile phones.
Oh didn’t see that. This is really cool! I suppose it does work similarly to hardware codecs, with the same very big trade-off being you get locked into a specific model. But considering this is an emerging technology maybe they could be made small enough to have multiple models on a single chip! (similar to codecs) Or just have one really big model that would be more future proof, and the price-performance would be so much greater than running the model on a GPU.
Oh didn’t see that. This is really cool! I suppose it does work similarly to hardware codecs, with the same very big trade-off being you get locked into a specific model. But considering this is an emerging technology maybe they could be made small enough to have multiple models on a single chip! (similar to codecs) Or just have one really big model that would be more future proof, and the price-performance would be so much greater than running the model on a GPU.