Why would AI want a continuous deal with Reddit? Don’t they get all the data they need the first time? I doubt the new content is worth as much as the previous deal… maybe I don’t understand what these deals are for.
Presuming they took all the data, a one-time deal would only be good if knowledge gathering actually stopped after the cutoff year- but for recent things like tech and news, the models have to keep learning and adding to their repositories.
The returns on that value sharply diminish, of course, but I think they’re still necessary. Which will leave everybody in a bind that is very funny.
Probably because Reddit has lawyers, and money, and a little willingness to lock down their content. Unlike individual creators, they can actually file a lawsuit
Why would AI want a continuous deal with Reddit? Don’t they get all the data they need the first time? I doubt the new content is worth as much as the previous deal… maybe I don’t understand what these deals are for.
Half the joke is that Reddit was ground zero for AI slop even before AI had gone mainstream.
The company got harvested back before the AI firms were overly worried with cross-contamination.
Presuming they took all the data, a one-time deal would only be good if knowledge gathering actually stopped after the cutoff year- but for recent things like tech and news, the models have to keep learning and adding to their repositories.
The returns on that value sharply diminish, of course, but I think they’re still necessary. Which will leave everybody in a bind that is very funny.
Why do they need a deal? Can’t they just steal it like the rest of their training data?
The API restrictions and login requirements are meant to make scraping hard enough to make a deal worthwhile.
Probably because Reddit has lawyers, and money, and a little willingness to lock down their content. Unlike individual creators, they can actually file a lawsuit
Other than that, probably it’s a licensing agreement that makes AI trainers keep paying if they still using that dataset.