Almost all these books were either headed to the dump or the recycling center.
My wife works at a law firm that contracts with “Books By The Yard”, which provides books purely for office aesthetics. For a few hundred bucks you can plaster a bookshelf full of material nobody will ever read, because they’re such a commodity.
It’s so crazy to see people ingesting this news from the conspiracy-brain perspective of “The AI companies are stealing all the knowledge!” without recognizing the more pressing reality of decade upon decade of publisher overproduction, resulting in a total devaluing of physical media for its academic importance.
Imagine going into hysterics because a warehouse full of shitty airport books went up in smoke, like it’s the Library of Alexandra that just burned down.
Are you sure they’re digitizing Book By the Yard quality material?
Consider this. Purchasing, shipping, and physically scanning a book is the single most expensive way for an AI company to acquire the text of a work. It’s been widely documented in court cases that these companies engage in mass IP theft. They’re literally running massive torrenting farms, grabbing copies of every film, song, book, etc. that they can get their hands on.
What kinds of texts are most likely to be found pirated on the internet? It’s the mass market stuff. The common stuff was digitized long ago. They can just download that. They can pirate it. They can buy the ebook. There’s no need for OpenAI to purchase and destructively scan the works of Steven King. No shade on the man, but his works aren’t exactly hard to find. I’m sure they can just find a torrent.
There’s little value in scanning the mass market books that are produced in enormous quantities. They probably don’t buy “Book by the Yard” books, because they already have digital copies of those.
But the rare, long out-of-print stuff? The stuff that you actually cannot find a legal or illegal copy of online? That’s only stuff worth paying money to buy, ship, and scan.
Doing anything physical, especially at scale, is slow and expensive. I would guess a good portion of these books, perhaps an outright majority, have simply never been digitized.
re you sure they’re digitizing Book By the Yard quality material?
They were more than happy to gobble up Reddit posts and Twitter farts and the worst of 4chan. You don’t need to ingest Ulysses exclusively to train a system.
Doing anything physical, especially at scale, is slow and expensive.
Which is why you get people paid pennies an hour to break the physical copies down and feed them into scanners.
would guess a good portion of these books, perhaps an outright majority, have simply never been digitized.
Possibly not. But that’s a sign it isn’t great literature or coveted texts.
You’re confusing uniqueness with quantity. 4chan posts are not Ulysses, but they are still useful data.
My point was not that they would refuse to scan Book By the Yard quality works because they would consider them unsuitable. My point was that they don’t need to scan mass market Book by the Yard stuff, as they already have digital copies of it.
And scanning books is not as cheap as you think it is. Even using labor in low income countries, the cost of scanning is still vastly greater than just downloading a file.
My wife works at a law firm that contracts with “Books By The Yard”, which provides books purely for office aesthetics. For a few hundred bucks you can plaster a bookshelf full of material nobody will ever read, because they’re such a commodity.
It’s so crazy to see people ingesting this news from the conspiracy-brain perspective of “The AI companies are stealing all the knowledge!” without recognizing the more pressing reality of decade upon decade of publisher overproduction, resulting in a total devaluing of physical media for its academic importance.
Imagine going into hysterics because a warehouse full of shitty airport books went up in smoke, like it’s the Library of Alexandra that just burned down.
Are you sure they’re digitizing Book By the Yard quality material?
Consider this. Purchasing, shipping, and physically scanning a book is the single most expensive way for an AI company to acquire the text of a work. It’s been widely documented in court cases that these companies engage in mass IP theft. They’re literally running massive torrenting farms, grabbing copies of every film, song, book, etc. that they can get their hands on.
What kinds of texts are most likely to be found pirated on the internet? It’s the mass market stuff. The common stuff was digitized long ago. They can just download that. They can pirate it. They can buy the ebook. There’s no need for OpenAI to purchase and destructively scan the works of Steven King. No shade on the man, but his works aren’t exactly hard to find. I’m sure they can just find a torrent.
There’s little value in scanning the mass market books that are produced in enormous quantities. They probably don’t buy “Book by the Yard” books, because they already have digital copies of those.
But the rare, long out-of-print stuff? The stuff that you actually cannot find a legal or illegal copy of online? That’s only stuff worth paying money to buy, ship, and scan.
Doing anything physical, especially at scale, is slow and expensive. I would guess a good portion of these books, perhaps an outright majority, have simply never been digitized.
They were more than happy to gobble up Reddit posts and Twitter farts and the worst of 4chan. You don’t need to ingest Ulysses exclusively to train a system.
Which is why you get people paid pennies an hour to break the physical copies down and feed them into scanners.
Possibly not. But that’s a sign it isn’t great literature or coveted texts.
You’re confusing uniqueness with quantity. 4chan posts are not Ulysses, but they are still useful data.
My point was not that they would refuse to scan Book By the Yard quality works because they would consider them unsuitable. My point was that they don’t need to scan mass market Book by the Yard stuff, as they already have digital copies of it.
And scanning books is not as cheap as you think it is. Even using labor in low income countries, the cost of scanning is still vastly greater than just downloading a file.
Solid point