cross-posted from: https://lemmy.dbzer0.com/post/73048134
Artificial intelligence labs are in a new arms race to buy up millions of rare books, slicing them open, scanning the pages and pulping the remains — sparking concerns that the last remaining copies of out-of-print texts are being destroyed on an industrial scale.
ISBNdb notes that “print books from the pre-LLM era are structurally guaranteed to be free of this contamination”.
“Millions of the most valuable books have never been digitised. They exist only in physical form, scattered across library shelves, used bookstores, and out-of-print catalogues. We get them to you at scale.”


Wow that’s really fucked up and completely unnecessary.
They could have just hired a bunch of broke students to flip the pages of the books or something but no they need the data yesterday gotta make the line go up
didn’t want any grubby human fingerprints on that sweet sweet data
It is the quickest and easiest way to digitize books. We used to cut the spines off of text books and scan them with our schools industrial photo copier to make the PDFs available to everyone.
It depends on what is meant by necessary, because all of this behavior is “made necessary” to maximize profit for the companies, but it’s extremely wasteful and makes the world worse than it otherwise quite easily could be, so I wouldn’t call it necessary on that level.
I just mean unnecessary in the sense that one doesn’t need to destroy a book to scan its contents
I understand, but if they weren’t doing this whole operation then they wouldn’t give a shit about scanning these books in the first place, which is why I said there are two types of necessity. They have every incentive to destroy the books to be more competitive. The fact that it’s worse for the rest of us is to them incidental to what is “necessitated” by the market.
Right, I understand. This happened to strike me as particularly heinous.