• artyom@piefed.social
    link
    fedilink
    English
    arrow-up
    40
    ·
    edit-2
    10 hours ago

    That’s true but they probably didn’t account for AI data scrapers ramfucking your server so they could steal all the value you assembled for general consumption and serve it themselves for profit.

    • grue@lemmy.world
      link
      fedilink
      English
      arrow-up
      41
      ·
      10 hours ago

      The scraping wouldn’t be a problem if Reddit simply provided an RSS feed or other data-efficient API. The “ramfucking” is caused by the attempt to block bots; it is entirely self-inflicted.

      Remember, it’s all our content to begin with and Reddit does not have any right to try to lock it up for itself.


      That doesn’t mean I like all the AI bullshit going on, BTW. But the problem is the generation of the slop, not the data accessibility.

      • artyom@piefed.social
        link
        fedilink
        English
        arrow-up
        33
        ·
        10 hours ago

        The scraping wouldn’t be a problem if Reddit simply provided an RSS feed or other data-efficient API

        That’s simply not true. These bots are essentially DDOSing the entire internet, API or not.

        • grue@lemmy.world
          link
          fedilink
          English
          arrow-up
          15
          ·
          9 hours ago

          Okay, if efficient APIs existed and they weren’t incompetently failing to use them, it wouldn’t be a problem. Happy now?

          (I should’ve addressed that in my previous comment, as I was aware of how one of the Lemmy instances was taken down by scrapers the other day despite the fact that they could easily get all the content simply by consuming ActivityPub directly. But I was naively hoping it wouldn’t be necessary because, as you can see from this text, it would’ve cluttered up my writing with double the words.)

      • so why was I getting hit with over 1,400,000 request a day to the web URI and not the API by some bot farm in China the other week. They were also hitting other lemmy instances.

        I blocked the fuckers, no qualms at all.

        Even if they were using the API they were not being nice about their shit.

      • Toga77@lemmy.world
        link
        fedilink
        English
        arrow-up
        6
        ·
        9 hours ago

        But we all know AI companies are unethically scraping and selling shit back to us right? I just really need people to acknowledge that.

        • grue@lemmy.world
          link
          fedilink
          English
          arrow-up
          2
          ·
          9 hours ago

          It is the “selling shit back to us” specifically, not the “scraping,” that’s the unethical part. If the AI companies were doing the same scraping (and destructive rare book scanning, for that matter), but were using the data to populate archive.org, would it still be a problem? I would argue “no.”

          • Tim_Bisley@piefed.social
            link
            fedilink
            English
            arrow-up
            2
            ·
            2 hours ago

            I think it would be a problem because the scrapers are hammering all types of websites from small forums to reddit with tens of thousands of unique ip addresses at a time. Websites that have neither the money, hardware, or protection had to figure out solutions really quick or suffer what is essentially a constant ddos attack. This is the reality of the web now, it’s just an incredibly hostile place.