• Riskable@programming.dev
    link
    fedilink
    English
    arrow-up
    16
    arrow-down
    33
    ·
    24 hours ago

    Is it really destroying though? They’re digitizing them, and publishers still have the digital copies ready to print more at any time. So it’s not like they’re destroying the texts, they’re just shifting them.

    Nobody complained when Google did this over a decade ago 🤷

    When you say they’re “destroying the books” you make it sound like they’re erasing one of the last known copy of some important work when in reality, most of these books were purchased in bulk from bookstores and libraries that were planning on discarding them anyway.

    Almost all these books were either headed to the dump or the recycling center. They’re just being digitized on the way.

    • Hawke@lemmy.world
      link
      fedilink
      English
      arrow-up
      57
      arrow-down
      1
      ·
      24 hours ago

      Digitized for private consumption.

      Nobody cares when Google did this or when archive.org does this, because they’re sharing the results with the world. (Idiotic shortsighted lawsuits from the authors guild notwithstanding)

      • antonim@lemmy.world
        link
        fedilink
        English
        arrow-up
        5
        ·
        edit-2
        16 hours ago

        Google (and/or the libraries that collaborate with Google Books) and Archive.org usually scan non-destructively. Some of Google’s book scans still show the fingers holding the corners of the pages, some of them were taken mistakenly mid-flip, etc. So they clearly used whole, normally bound books.

      • SkaveRat@discuss.tchncs.de
        link
        fedilink
        English
        arrow-up
        12
        ·
        23 hours ago

        in this case they often are actually destroying them. they take them apart, beause it’s easier to scan than to use a proper book scanner

        • Hawke@lemmy.world
          link
          fedilink
          English
          arrow-up
          5
          arrow-down
          1
          ·
          22 hours ago

          Yeah that’s true of Google as well I believe. I think Archive is more careful, but I’m sure some books get damaged in the process there as well.

          • Zarobi@aussie.zone
            link
            fedilink
            English
            arrow-up
            4
            ·
            12 hours ago

            Having done a lot of scanning in the past, no matter how gentle you are, the book will be damaged; at least a little bit. Even if you do it manually, by hand, really slowly.

            It’s the unfortunate reality, because books are not designed to be held open and pressed flat onto an unyielding surface. If you don’t press flat, you’ll get curved pages and bad quality (dark) images, which is sometimes ok and fixable by software, but often not. You can get really fancy expensive machines, but they still don’t fully solve the problem.

            I did scanlation editing stuff for a while, and half my job was making the raw page scans look presentable, because the scanners didn’t want to destroy their manga (understandable).

            https://en.wikipedia.org/wiki/Book_scanning#Methods

            • smh@slrpnk.net
              link
              fedilink
              English
              arrow-up
              3
              ·
              12 hours ago

              This matches my experience. My college archive digitizes a few yearbooks a year (for whatever classes are having as big reunion) and it’s a pain. For modern yearbooks (70s and later) we have a dozen or more copies, so we take one apart and feed it through the feed scanner, which works well. We then store the loose pages in a folder, in case we need to rescan anything.

              Older yearbooks (where we don’t have so many copies) we use the flatbed scanner and the pages come out warped, so we hand fix each page in Photoshop.

              • Zarobi@aussie.zone
                link
                fedilink
                English
                arrow-up
                1
                ·
                11 hours ago

                Yepp… memories of using the light level curve tool and white point and black point and the perspective tool. Then going in with the white brush and deleting any specks of dust that remained, and comparing to the original to make sure no lines got deleted

    • UnderpantsWeevil@lemmy.world
      link
      fedilink
      English
      arrow-up
      15
      ·
      edit-2
      20 hours ago

      Almost all these books were either headed to the dump or the recycling center.

      My wife works at a law firm that contracts with “Books By The Yard”, which provides books purely for office aesthetics. For a few hundred bucks you can plaster a bookshelf full of material nobody will ever read, because they’re such a commodity.

      It’s so crazy to see people ingesting this news from the conspiracy-brain perspective of “The AI companies are stealing all the knowledge!” without recognizing the more pressing reality of decade upon decade of publisher overproduction, resulting in a total devaluing of physical media for its academic importance.

      Imagine going into hysterics because a warehouse full of shitty airport books went up in smoke, like it’s the Library of Alexandra that just burned down.

      • isleepinahammock@lemmy.blahaj.zone
        link
        fedilink
        English
        arrow-up
        4
        ·
        12 hours ago

        Are you sure they’re digitizing Book By the Yard quality material?

        Consider this. Purchasing, shipping, and physically scanning a book is the single most expensive way for an AI company to acquire the text of a work. It’s been widely documented in court cases that these companies engage in mass IP theft. They’re literally running massive torrenting farms, grabbing copies of every film, song, book, etc. that they can get their hands on.

        What kinds of texts are most likely to be found pirated on the internet? It’s the mass market stuff. The common stuff was digitized long ago. They can just download that. They can pirate it. They can buy the ebook. There’s no need for OpenAI to purchase and destructively scan the works of Steven King. No shade on the man, but his works aren’t exactly hard to find. I’m sure they can just find a torrent.

        There’s little value in scanning the mass market books that are produced in enormous quantities. They probably don’t buy “Book by the Yard” books, because they already have digital copies of those.

        But the rare, long out-of-print stuff? The stuff that you actually cannot find a legal or illegal copy of online? That’s only stuff worth paying money to buy, ship, and scan.

        Doing anything physical, especially at scale, is slow and expensive. I would guess a good portion of these books, perhaps an outright majority, have simply never been digitized.

        • UnderpantsWeevil@lemmy.world
          link
          fedilink
          English
          arrow-up
          1
          ·
          2 hours ago

          re you sure they’re digitizing Book By the Yard quality material?

          They were more than happy to gobble up Reddit posts and Twitter farts and the worst of 4chan. You don’t need to ingest Ulysses exclusively to train a system.

          Doing anything physical, especially at scale, is slow and expensive.

          Which is why you get people paid pennies an hour to break the physical copies down and feed them into scanners.

          would guess a good portion of these books, perhaps an outright majority, have simply never been digitized.

          Possibly not. But that’s a sign it isn’t great literature or coveted texts.

          • isleepinahammock@lemmy.blahaj.zone
            link
            fedilink
            English
            arrow-up
            1
            ·
            37 minutes ago

            You’re confusing uniqueness with quantity. 4chan posts are not Ulysses, but they are still useful data.

            My point was not that they would refuse to scan Book By the Yard quality works because they would consider them unsuitable. My point was that they don’t need to scan mass market Book by the Yard stuff, as they already have digital copies of it.

            And scanning books is not as cheap as you think it is. Even using labor in low income countries, the cost of scanning is still vastly greater than just downloading a file.

    • minus@lemmy.zip
      link
      fedilink
      English
      arrow-up
      23
      arrow-down
      1
      ·
      edit-2
      22 hours ago

      When they destroy the physical copy they remove them from the antique stores market which often rely on circulation.

      Edit: Also physical copies don’t require electricity, a device and internet access plus they are something you can own.

      • Riskable@programming.dev
        link
        fedilink
        English
        arrow-up
        6
        ·
        edit-2
        18 hours ago

        If you think any more than 0.1% of these physical books would ever have ended up in antique bookstores, you’re dreaming.

        Think about how many books out there are things like Donald Trump’s biography, or pointless drivel from non-experts, self-help books that tell people to down “essential oils”, old editions of programming books, or just plain shitty fiction that never sold much in the first place.

        It’s ok to throw trash away! Really!

        • isleepinahammock@lemmy.blahaj.zone
          link
          fedilink
          English
          arrow-up
          2
          ·
          12 hours ago

          Why would they scan that stuff though? That kind of mass market (human made) slop is easily available in digital form. We know the AI companies have engaged in mass digital piracy, including running massive torrenting operations. So if a digital copy exists, they probably already have it. And even if they have to buy it, purchasing an ebook is a lot cheaper than buying a physical one, shipping it, paying someone to scan it, etc.

          I would think the old, the out-of-print, the rare, and never-before digitized are the only things worth buying and physically scanning in 2026. Everything else has already been scanned or was born digital-native.

      • XLE@piefed.social
        link
        fedilink
        English
        arrow-up
        3
        arrow-down
        1
        ·
        22 hours ago

        At least digital storage prices have only been going down in recent months, right? /s

      • Snot Flickerman@lemmy.blahaj.zone
        link
        fedilink
        English
        arrow-up
        24
        arrow-down
        1
        ·
        23 hours ago

        The knowledge is retained and owned by a private corporation who now does not have to share what may have been a still under copyright, but now the corporation owns it? Make it make sense.

        • Riskable@programming.dev
          link
          fedilink
          English
          arrow-up
          2
          ·
          19 hours ago

          Yes. That does make sense.

          If you bought a book, scanned it—destroying it in the process—then read it on your computer, that would be completely acceptable.

          Why is it wrong when a corporation does the same thing?

          They’re not claiming ownership of the copyrights, just ownership of a copy. Which is how copyright works.

          • BlaestEgnen@feddit.dk
            link
            fedilink
            English
            arrow-up
            1
            ·
            10 hours ago

            Now we just need to proof they’ve copied the files onto a second hard drive, so they own two copies

            If we’re gonna keep it as a copyright issue, I think people are mostly mad they’re not able to reprint the books themselves if they wanted to. We’re not speaking from a copyright position