• FlashMobOfOne@lemmy.world
    link
    fedilink
    English
    arrow-up
    1
    ·
    5 分钟前

    Yeah, pretty awful on so many different levels.

    Brilliant decision to train their models on the worst content to begin with.

  • Deacon@lemmy.world
    link
    fedilink
    English
    arrow-up
    4
    ·
    47 分钟前

    You can only stave off model collapse so long.

    It’s only a matter of time before most of the data going in to LLMs came out of one too.

    Like I’m not even good at math or anything, much less computer science, but that seems obvious and inevitable to me.

    Maybe someone smarter than me can chime in.

  • meme_historian@lemmy.dbzer0.com
    link
    fedilink
    English
    arrow-up
    157
    ·
    6 小时前

    Bruh, this is literally like scientists scavenging metal from shipwrecks that predate the era of nuclear bomb testing, cause they need steel that isn’t contaminated by nuclear fallout for certain measuring equipment 😵‍💫😵‍💫

    Great analogy for the shit we’re currently saturating our digital world with

    • Courtney (she/her/they) @lemmy.blahaj.zone
      link
      fedilink
      English
      arrow-up
      6
      ·
      1 小时前

      A guy I used to know who would go to estate sales looking for blacksmithing equipment once told me the market for old metalworking stuff is disappearing because too much stuff is being bought by companies én masse for this reason.

      I don’t know if I believe that companies are buying small lot auctions and estate sales for old anvils, but he was sure convinced.

      I like to joke I’m sitting on a pile of cash because my anvil isn’t radioactive. (it’s not a very good joke, but it carries some weight!)

    • AbouBenAdhem@lemmy.world
      link
      fedilink
      English
      arrow-up
      7
      ·
      2 小时前

      Reminds me of Judge Holden in Blood Meridian, who sketched ancient artifacts in his personal notebook and then destroyed the originals.

      • Omgpwnies@lemmy.world
        link
        fedilink
        English
        arrow-up
        5
        ·
        1 小时前

        More that they are destroying the books in the process of digitizing them, they basically cut the pages out of the book and feed them into a scanner with a document feeder attached.

        • T156@lemmy.world
          link
          fedilink
          English
          arrow-up
          2
          ·
          23 分钟前

          It’s worth pointing out that this is standard practice when digitising books in most places, if they’re not irreplaceable. Mass-produced books would fall under that category.

          Turning the page and photographing what’s there is really only done for some books, which you can’t afford to destroy.

    • crusa187@lemmy.ml
      link
      fedilink
      English
      arrow-up
      31
      ·
      5 小时前

      We also have no guarantee these books will be available at a later date in these digitized formats, nor that they won’t be altered to fit certain narratives which is of course trivial to do with digital copies. At a minimum these need public verifiable checksums to be relied upon.

      I don’t like this and don’t trust it one bit.

      These AI companies already showed their hand in trying to corner the PC component market with an end goal of users no longer being able to own their own PCs. Now they’re potentially trying to do this with all written recorded information.

      How long until Altman’s Firemen show up at your door to burn your books?

    • Riskable@programming.dev
      link
      fedilink
      English
      arrow-up
      12
      ·
      6 小时前

      Is it really destroying though? They’re digitizing them, and publishers still have the digital copies ready to print more at any time. So it’s not like they’re destroying the texts, they’re just shifting them.

      Nobody complained when Google did this over a decade ago 🤷

      When you say they’re “destroying the books” you make it sound like they’re erasing one of the last known copy of some important work when in reality, most of these books were purchased in bulk from bookstores and libraries that were planning on discarding them anyway.

      Almost all these books were either headed to the dump or the recycling center. They’re just being digitized on the way.

      • UnderpantsWeevil@lemmy.world
        link
        fedilink
        English
        arrow-up
        12
        ·
        edit-2
        2 小时前

        Almost all these books were either headed to the dump or the recycling center.

        My wife works at a law firm that contracts with “Books By The Yard”, which provides books purely for office aesthetics. For a few hundred bucks you can plaster a bookshelf full of material nobody will ever read, because they’re such a commodity.

        It’s so crazy to see people ingesting this news from the conspiracy-brain perspective of “The AI companies are stealing all the knowledge!” without recognizing the more pressing reality of decade upon decade of publisher overproduction, resulting in a total devaluing of physical media for its academic importance.

        Imagine going into hysterics because a warehouse full of shitty airport books went up in smoke, like it’s the Library of Alexandra that just burned down.

      • minus@lemmy.zip
        link
        fedilink
        English
        arrow-up
        20
        ·
        edit-2
        4 小时前

        When they destroy the physical copy they remove them from the antique stores market which often rely on circulation.

        Edit: Also physical copies don’t require electricity, a device and internet access plus they are something you can own.

        • Riskable@programming.dev
          link
          fedilink
          English
          arrow-up
          2
          ·
          edit-2
          39 分钟前

          If you think any more than 0.1% of these physical books would ever have ended up in antique bookstores, you’re dreaming.

          Think about how many books out there are things like Donald Trump’s biography, or pointless drivel from non-experts, self-help books that tell people to down “essential oils”, old editions of programming books, or just plain shitty fiction that never sold much in the first place.

          It’s ok to throw trash away! Really!

          • xthexder@l.sw0.com
            link
            fedilink
            English
            arrow-up
            2
            ·
            1 小时前

            I guess you must live in that part of Alaska where it’s night time for half the year and don’t have daylight.

        • XLE@piefed.social
          link
          fedilink
          English
          arrow-up
          3
          ·
          4 小时前

          At least digital storage prices have only been going down in recent months, right? /s

      • Hawke@lemmy.world
        link
        fedilink
        English
        arrow-up
        49
        ·
        6 小时前

        Digitized for private consumption.

        Nobody cares when Google did this or when archive.org does this, because they’re sharing the results with the world. (Idiotic shortsighted lawsuits from the authors guild notwithstanding)

        • SkaveRat@discuss.tchncs.de
          link
          fedilink
          English
          arrow-up
          10
          ·
          6 小时前

          in this case they often are actually destroying them. they take them apart, beause it’s easier to scan than to use a proper book scanner

          • Hawke@lemmy.world
            link
            fedilink
            English
            arrow-up
            3
            ·
            4 小时前

            Yeah that’s true of Google as well I believe. I think Archive is more careful, but I’m sure some books get damaged in the process there as well.

        • Snot Flickerman@lemmy.blahaj.zone
          link
          fedilink
          English
          arrow-up
          22
          ·
          5 小时前

          The knowledge is retained and owned by a private corporation who now does not have to share what may have been a still under copyright, but now the corporation owns it? Make it make sense.

          • Riskable@programming.dev
            link
            fedilink
            English
            arrow-up
            1
            ·
            46 分钟前

            Yes. That does make sense.

            If you bought a book, scanned it—destroying it in the process—then read it on your computer, that would be completely acceptable.

            Why is it wrong when a corporation does the same thing?

            They’re not claiming ownership of the copyrights, just ownership of a copy. Which is how copyright works.

  • kescusay@lemmy.world
    link
    fedilink
    English
    arrow-up
    34
    ·
    6 小时前

    I’m waiting for the day when desperate LLM companies start paying people to post real human content, only for those people to just ask ChatGPT to do it.

    • veryblandusername@fedinsfw.app
      link
      fedilink
      English
      arrow-up
      3
      ·
      2 小时前

      This is already happening.

      They are paying people bottom market rates to generate unique written content and people are just having AI write it for them.

    • pdxfed@lemmy.world
      link
      fedilink
      English
      arrow-up
      9
      ·
      6 小时前

      “now you only posted 4 times today Jenny, do you just want to do the minimum?”

    • Riskable@programming.dev
      link
      fedilink
      English
      arrow-up
      3
      ·
      6 小时前

      Big AI mostly switched to synthetic training data anyway. The books they’re digitizing are being used to gather knowledge, not writing styles or logic (mostly).

      As in, when you ask ChatGPT how long some book is, it can just go check (if it’s in the database). It’s also useful if you ask about that book or about knowledge contained in that book. It’ll even reference books now (if you demand that in your prompt).

      It’s not the same as earlier LLM tech which relied on scanned text to figure out how to respond to any given prompt (from a language standpoint). The “language” part of LLMs is a solved problem now (thanks to the synthetic training). At least for English 🤷

  • SparroHawc@piefed.world
    link
    fedilink
    English
    arrow-up
    14
    ·
    6 小时前

    Nice that they’re actually buying them this time instead of just downloading torrents of book collections again.