• isleepinahammock@lemmy.blahaj.zone
    link
    fedilink
    English
    arrow-up
    4
    ·
    3 days ago

    Why would they scan that stuff though? That kind of mass market (human made) slop is easily available in digital form. We know the AI companies have engaged in mass digital piracy, including running massive torrenting operations. So if a digital copy exists, they probably already have it. And even if they have to buy it, purchasing an ebook is a lot cheaper than buying a physical one, shipping it, paying someone to scan it, etc.

    I would think the old, the out-of-print, the rare, and never-before digitized are the only things worth buying and physically scanning in 2026. Everything else has already been scanned or was born digital-native.

    • Riskable@programming.dev
      link
      fedilink
      English
      arrow-up
      1
      ·
      22 hours ago

      Legal reasons: When you “purchase” an ebook you’re actually just licensing it and nearly all ebook licenses exclude the ability to do anything with the ebook other than read it yourself.

      They could get a commercial license to get big ebook libraries like Anthropic did, but they can’t, really, because Anthropic paid for an exclusive license. Which means that if other AI companies want to compete, they sort of have to buy books in bulk and tear them apart to scan them.

      • isleepinahammock@lemmy.blahaj.zone
        link
        fedilink
        English
        arrow-up
        1
        ·
        22 hours ago

        True, but irrelevant. Why would they care at all about the terms of a license? Again, they’ll happily engage in outright mass torrenting. Their legal theory is that using works to train LLMs is simply fair use. The ebook sellers may disagree, but it won’t stop them.