All of their business decisions for the past few years have been centered around monetizing Reddit content as training data for machine learning models. Most likely this recent decision is an attempt to block data scraping.
Even within the lemmy user community people still occasionally refer to reddit because there isn’t a better source for some things. Outside of lemmy, reddit is a household name alongside facebook, snapchat &etc; people who don’t use it still know what it is. I think your conclusion is more aspirational than currently real.
Oh absolutely, I think that is already happening, but it hasn’t reached broad public awareness yet. Reddit has 20 years of momentum, it won’t stop all at once.
I suspect the company is aware that user engagement is dropping, along with the quality of new content. Typical of most corporations, their response is to play defensive, to try to maximize the value of what they already have, to scratch and claw and squeeze every drop from the existing data and userbase. This will accelerate the decline, obviously, but when a company is in this state they consider any strategy besides protectionism to be too high a risk.
They will ruin themselves with cowardice, to be sure, but they’re still coasting on momentum for the moment. They will continue making the same types of decisions, at an accelerating pace, as the desperation sets in.
Volume, and breadth of subject material. For language model purposes the specific content isn’t really important, it’s the wide variety of language samples from many sources, all in the same data format. You could get similar samples from other platforms, but you’d have to compile them from multiple sources and then standardize them somehow for input as training data.
All of their business decisions for the past few years have been centered around monetizing Reddit content as training data for machine learning models. Most likely this recent decision is an attempt to block data scraping.
It’s gonna block sign ups. Why the fuck do I sign up without seeing shit?
Not like reddit is relevant these days.
Even within the lemmy user community people still occasionally refer to reddit because there isn’t a better source for some things. Outside of lemmy, reddit is a household name alongside facebook, snapchat &etc; people who don’t use it still know what it is. I think your conclusion is more aspirational than currently real.
The opacity is going to make the issue very real. A site like that goes stagnant without new blood.
It doesn’t need new blood. It needs infinite bots creating infinite content with no way to prove whether or not they’re bots or humans.
Wouldn’t you know it they’ve got it covered.
Oh absolutely, I think that is already happening, but it hasn’t reached broad public awareness yet. Reddit has 20 years of momentum, it won’t stop all at once.
I suspect the company is aware that user engagement is dropping, along with the quality of new content. Typical of most corporations, their response is to play defensive, to try to maximize the value of what they already have, to scratch and claw and squeeze every drop from the existing data and userbase. This will accelerate the decline, obviously, but when a company is in this state they consider any strategy besides protectionism to be too high a risk.
They will ruin themselves with cowardice, to be sure, but they’re still coasting on momentum for the moment. They will continue making the same types of decisions, at an accelerating pace, as the desperation sets in.
Everything posted in the last few years should be considered poison anyway, so I don’t know why anyone would pay for it to use as training data.
Why anyone would train an AI on Reddit is beyond me, the only worse platform I can think of is 4chan.
Because it has years of helpful troubleshooting guides, solutions to problems, and conversations. also there is an llm trained off 4chan, gpt-4chan.
Volume, and breadth of subject material. For language model purposes the specific content isn’t really important, it’s the wide variety of language samples from many sources, all in the same data format. You could get similar samples from other platforms, but you’d have to compile them from multiple sources and then standardize them somehow for input as training data.