Would be a terrible shame if lots of people opted out.

If you are a EU citizen you might also want to write a complaint to privacy@huggingface.co because they are collecting your personally identifiable information in machine-readable form which they are distributing to third parties.

  • SpaceCowboy@lemmy.ca
    link
    fedilink
    arrow-up
    2
    ·
    40 minutes ago

    A decade ago, if someone asked someone working on an Open Source project if they’d be ok with an AI reading their code and learning from it, they’d most likely say “yeah that sounds really cool!”

    Somehow the tech-bros have fucked up AI so much that something that should be really cool seems creepy, lame, and nefarious all at once.

  • katy ✨@piefed.blahaj.zone
    link
    fedilink
    English
    arrow-up
    56
    ·
    4 hours ago

    We want to give developers agency over their source code by letting them decide whether or not it should be used to develop and evaluate machine learning models.

    fuck them seriously; if you want to do that then don’t steal the repositories in the first place.

    • lps2@lemmy.ml
      link
      fedilink
      arrow-up
      15
      ·
      2 hours ago

      Kinda hope it uses my code. It’s so terrible there’s no doubt it will make the resulting code from the model worse even if the impact is miniscule

    • kibiz0r@midwest.social
      link
      fedilink
      English
      arrow-up
      32
      ·
      3 hours ago

      Meanwhile at work we just had a training course that specifically said doing “opt out” instead of “opt in” violates the principle of informed consent.

    • JakenVeina@midwest.social
      link
      fedilink
      arrow-up
      5
      ·
      2 hours ago

      Unless I’m mistaken, this wasn’t written by the folks that scaped GitHub in the first place, someone just wrote a small tool to semi-automate the process of searching the scraped dats, and submitting a GitHub issue to have it removed.

  • goatbeard@beehaw.org
    link
    fedilink
    arrow-up
    1
    ·
    edit-2
    48 minutes ago

    Since they stole my paper on ethics in computer science, maybe the model will learn to act better than its owners

  • chicken@lemmy.dbzer0.com
    link
    fedilink
    arrow-up
    15
    ·
    3 hours ago

    A little bit infuriating since huggingface itself requires login to access a large portion of the content on their site

    • DarkCloud@lemmy.world
      link
      fedilink
      arrow-up
      3
      ·
      edit-2
      3 hours ago

      When I back something up, I save a copy and add “backup” to the name… Because I’m advanced.

      If I’m feeling really good and healthy, I’ll even put it on a usb stick.

    • Eager Eagle@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      ·
      edit-2
      3 hours ago

      note that former users would have needed to remove their GitHub data before August 2025 to not be in this dataset

    • korendian@piefed.social
      link
      fedilink
      English
      arrow-up
      1
      ·
      10 minutes ago

      This is what confuses me. The internet archive has been archive the entire Internet for years. Yet AI does the same to make a way for people to code easier and it is a problem all of the sudden?

    • cecilkorik@lemmy.ca
      link
      fedilink
      English
      arrow-up
      2
      ·
      edit-2
      30 minutes ago

      As long as the datasets are open, it is our best hope. I know it doesn’t compensate the people whose work’s copyright and licenses have been violated, but I think it’s the only realistic hope we’ve got of getting out of this informational dystopia with a reasonably intact library of humanity’s knowledge that hasn’t been locked down and/or monetized. The AI scrapers and generators are in the process of burning down the great library of Alexandria that the Internet had become, and we are already starting to feel its loss. We cannot stop the wave of toxic pollution that is spreading through all our digital content now, but the archives from before this apocalypse started will become the most valuable thing humanity has ever produced. This is information war, and we are losing.

    • lavember@programming.dev
      link
      fedilink
      arrow-up
      3
      ·
      2 hours ago

      Sure. Let’s see if they use it for endeavours in the same spirit.

      Or are you fine with they using this data for-profit without benefitting the public by also making it open?

      (not talking about hf, I know starcoder. just in general)

    • kibiz0r@midwest.social
      link
      fedilink
      English
      arrow-up
      10
      ·
      3 hours ago

      I did it to invite collaboration and connect with other developers with similar interests. FOSS is more about building communities than building software, after all.

      I did not anticipate that it could be (legally) used to dismantle the kinds of communities I wanted to build. (I did anticipate that it could be illegally used to that end, but historically that has tended to cause a Streisand Effect, so that risk seemed worth it.)