Artificial intelligence labs are in a new arms race to buy up millions of rare books, slicing them open, scanning the pages and pulping the remains — sparking concerns that the last remaining copies of out-of-print texts are being destroyed on an industrial scale.

ISBNdb notes that “print books from the pre-LLM era are structurally guaranteed to be free of this contamination”.

“Millions of the most valuable books have never been digitised. They exist only in physical form, scattered across library shelves, used bookstores, and out-of-print catalogues. We get them to you at scale.”

  • Goodlucksil@lemmy.dbzer0.com
    link
    fedilink
    English
    arrow-up
    31
    ·
    edit-2
    15 hours ago

    What’s the point of destroying books after scanning them instead of reselling them apart from *hurr durr we are evil*?

    • Jason2357@lemmy.ca
      link
      fedilink
      English
      arrow-up
      13
      ·
      9 hours ago

      Easier to scan if you cut the binding off, and it costs money to re-bind a book you purchased for 10 cents as part of a lot.

    • chiliedogg@lemmy.world
      link
      fedilink
      English
      arrow-up
      20
      ·
      12 hours ago

      Have you ever tried scanning a book? It’s a huge PITA.

      They just cut the binding off so they can be run through a document feeder.

    • leds@feddit.dk
      link
      fedilink
      English
      arrow-up
      7
      ·
      10 hours ago

      Do make sure that knowledge is only available through their model and otherwise lost to humanity.

      Same with bombarding small websites until they give up and pull the plug.

      Everything is fucked

        • daggermoon@lemmy.world
          link
          fedilink
          English
          arrow-up
          2
          ·
          12 hours ago

          Sorry, but I did an oopsie. I watched the video linked, they destroy the books to scan them. It’s really sad. Though I’m sure they would destroy them anyway.

          • Greyghoster@aussie.zone
            link
            fedilink
            English
            arrow-up
            1
            ·
            12 hours ago

            I assumed the process destroyed the book but keeping the contents from the competitors would definitely be seen as a business goal.

    • General_Effort@lemmy.world
      link
      fedilink
      English
      arrow-up
      8
      ·
      12 hours ago

      Copyright. They mustn’t make a copy. There is legal precedent that confirms it’s okay to transform the copy you bought into digital format. The Internet Archive relies a lot on that. I think they actually litigated it in the first place. So that’s why the copyright heads are going so absolutely apeshit. If no one’s charging you rent for using some data, then it’s “unethical”.

      • Jason2357@lemmy.ca
        link
        fedilink
        English
        arrow-up
        7
        ·
        edit-2
        9 hours ago

        The Internet Archive absolutely does not destroy books. They scan them the hard way with the binding still intact.

        • General_Effort@lemmy.world
          link
          fedilink
          English
          arrow-up
          1
          ·
          4 hours ago

          I was misremembering. Google won the precedent, when they were suing over Google Books. IA had only 1 big lawsuit and were forced to settle. The non-destructive scanning is dicey.

    • altkey (he\him)@lemmy.dbzer0.com
      link
      fedilink
      English
      arrow-up
      4
      ·
      14 hours ago

      Fair use loophole, they claim they convert one physical book into one digital with no illicit copies, while also supporting the delusion of AI being an equivalent to a human reader. It is weird on so many levels but no one of noticeable weight asked them wtf.