I have used crowd collected data for research. You have to curate each picture carefully and cannot trust labels, there are many erroneous identifications.
Maybe if you have access to pre-curated you can trust it a bit, but if you are a researcher and use images from these websites without curation you are already doing it wrong…
AI altered images doesn’t change things too much, it probably just increases the numbers of errors that already existed…
The “failure mode” of AI editing is different though.
Humans (I guess) might mislabel something or take a bad shot. If they try to touch it up “traditionally” they could mess up the coloration at most.
But with AI editing, now you have to watch out for fine details you’d normally use for identification being completely, convincingly fabricated, as the article points out, with altruistic intent from the user (who’s just trying to submit data that looks alright)
The solution is global AI literacy; but that’s not going so well.
For “AI editing”, I don’t have much of a problem if a model was trained to tweak the settings in something like Darktable to achieve a good baseline to start with. However, for contributing to research like this, I draw the line when models start generating their own pixels and overwriting the original image.
Hopefully iNaturalist and other similar groups start to look at the metadata of submitted images to help warn/educate end users about this problem. That would at least help with the AI literacy issue. Those metadata tags are already being placed there by the most popularly used tools.
on one hand it increases the percentage of invalid/unusable images. A mis-labeled image is still potentially useful, an image of a toucan in the Arctic isn’t. This means that even large datasets become largely unusable because of the volume of fake images.
on the other hand any AI picture that slips through, taints the dataset and results in expensive manual cleanup to ensure data is reliable.
I have used crowd collected data for research. You have to curate each picture carefully and cannot trust labels, there are many erroneous identifications.
Maybe if you have access to pre-curated you can trust it a bit, but if you are a researcher and use images from these websites without curation you are already doing it wrong…
AI altered images doesn’t change things too much, it probably just increases the numbers of errors that already existed…
The “failure mode” of AI editing is different though.
Humans (I guess) might mislabel something or take a bad shot. If they try to touch it up “traditionally” they could mess up the coloration at most.
But with AI editing, now you have to watch out for fine details you’d normally use for identification being completely, convincingly fabricated, as the article points out, with altruistic intent from the user (who’s just trying to submit data that looks alright)
The solution is global AI literacy; but that’s not going so well.
For “AI editing”, I don’t have much of a problem if a model was trained to tweak the settings in something like Darktable to achieve a good baseline to start with. However, for contributing to research like this, I draw the line when models start generating their own pixels and overwriting the original image.
Hopefully iNaturalist and other similar groups start to look at the metadata of submitted images to help warn/educate end users about this problem. That would at least help with the AI literacy issue. Those metadata tags are already being placed there by the most popularly used tools.
Yes, actually that would be great!
All this has happened so fast; it takes time to react I suppose.
I’d argue that AI does change things: