• iceberg314@slrpnk.net
    link
    fedilink
    English
    arrow-up
    23
    ·
    17 hours ago

    I’m a big fan of local AI and I think it has to be the future.

    It’s ridiculous l, like Bonsia AI’s Q1 models are like 3.5GB easily doing basic tasks that most people are asking 600GB flagship models.

    Who on earth would pay for something that needs a basically terabyte or RAM that only performs 10% better

    • eicker@lemmy.worldOP
      link
      fedilink
      English
      arrow-up
      14
      ·
      17 hours ago

      The industry keeps benchmarking against other labs instead of against user needs: If a 3.5GB model answers 95% of everyday questions well enough, the remaining few percent has to justify hundreds of gigabytes of weights, huge energy bills and constant cloud costs.

    • partofthevoice@lemmy.zip
      link
      fedilink
      English
      arrow-up
      4
      ·
      edit-2
      13 hours ago

      It’ll be “local AI” when they let me point to my own self hosted inference servers, rather than simply OpenAI, Anthropic, or Gemini as providers. Soon to include Apple provider, I guess.

      I’ve used the Apple Intelligence ecosystem. The models are slightly acceptable in extremely small context windows, but they completely shit the bed for any kind of practical ad-hoc use. Even if you try to dumb it down to like 6 words. It’s trash. It couldn’t even find a picture of my finger with the keyword “finger.” It couldn’t explain basic details of my phone… it was like interacting with a shittier version of ChatGPTs first release — much shittier.

      It told me I have an iPhone 18. I have a 17 Max Pro, the 18 hasn’t been released yet. I don’t plan to ever buy the 18, given the hardware regression with their charging port. But… as I said, it’s not even released yet.

      That’s fine with me, actually. If that’s all my phone can handle then so be it. But, when I eventually and obviously will want something more practical, my only options shouldn’t be to pay frontier cloud models if I want deep integration with my phone. The only alternative shouldn’t be to subscribe to a higher tier iCloud+.

      I can put a vpn on my phone to access vLLM or ollama locally. Why won’t they let me use that?

      • iturnedintoanewt@lemmy.world
        link
        fedilink
        English
        arrow-up
        2
        ·
        12 hours ago

        On android you can just use something like jegly Box and just run Gemma 4 or some other model. Decent results for an offline ai.