Artificial intelligence labs are in a new arms race to buy up millions of rare books, slicing them open, scanning the pages and pulping the remains — sparking concerns that the last remaining copies of out-of-print texts are being destroyed on an industrial scale.
ISBNdb notes that “print books from the pre-LLM era are structurally guaranteed to be free of this contamination”.
“Millions of the most valuable books have never been digitised. They exist only in physical form, scattered across library shelves, used bookstores, and out-of-print catalogues. We get them to you at scale.”


To be honest, I’d be MUCH less against this practice if this shredding at the very least included saving the original scans at at least 300 dpi for everyone to access, freely.
Ideally, they wouldn’t cut them by the spine, but even giving the despined ones to an archive/library would be just above the unpermissible line.
Despining, scanning and public access archiving is borderline permissible.
Also, if every AI company does this, then it’s more books gone. If they cooperated they’d only need to do it once. Good scans of all books in the public domain would be so useful for us all as society.