Presuming they took all the data, a one-time deal would only be good if knowledge gathering actually stopped after the cutoff year- but for recent things like tech and news, the models have to keep learning and adding to their repositories.
The returns on that value sharply diminish, of course, but I think they’re still necessary. Which will leave everybody in a bind that is very funny.
Probably because Reddit has lawyers, and money, and a little willingness to lock down their content. Unlike individual creators, they can actually file a lawsuit
Presuming they took all the data, a one-time deal would only be good if knowledge gathering actually stopped after the cutoff year- but for recent things like tech and news, the models have to keep learning and adding to their repositories.
The returns on that value sharply diminish, of course, but I think they’re still necessary. Which will leave everybody in a bind that is very funny.
Why do they need a deal? Can’t they just steal it like the rest of their training data?
The API restrictions and login requirements are meant to make scraping hard enough to make a deal worthwhile.
Probably because Reddit has lawyers, and money, and a little willingness to lock down their content. Unlike individual creators, they can actually file a lawsuit
Other than that, probably it’s a licensing agreement that makes AI trainers keep paying if they still using that dataset.