• auntieclokwise@lemmy.world
      link
      fedilink
      English
      arrow-up
      11
      arrow-down
      1
      ·
      10 hours ago

      Not to defend the AI crowd, but ever go to a Goodwill Outlet store? They usually have at least 1 gaylord box (the giant box on a pallet watermelons come in) full of books. Not sure how often they cycle those out, but I’m guessing pretty regularly. Multiply that by every semi major city Goodwill has a presence in. Add in other thrift sores with a similar quantity of books. You’re talking ALOT of books. SO MANY books are thrown out everyday. Including by libraries. I don’t really have a problem if some of those get diverted for a bit to be scanned. Yes, it’s sad to see books getting destroyed. But there’s alot of stuff people just really don’t want. And, if you REALLY want a villain, maybe you can blame copyright law? That’s part of the reason AI companies destroy the books - it helps to make their scanning efforts more legal. Absent that, they could probably be convinced to spend just a little more for nondestructive scanning, assuming they could find a home for the books they scanned (the sellers of these things couldn’t, so good luck).

      • GalacticRobot@lemmy.world
        link
        fedilink
        English
        arrow-up
        2
        ·
        3 hours ago

        I mean even when you go to Goodwill, it’s not like 99% of the books are rare or irreplaceable. How many copies of Patterson books does the world need to begin with?

  • Deacon@lemmy.world
    link
    fedilink
    English
    arrow-up
    37
    ·
    19 hours ago

    You can only stave off model collapse so long.

    It’s only a matter of time before most of the data going in to LLMs came out of one too.

    Like I’m not even good at math or anything, much less computer science, but that seems obvious and inevitable to me.

    Maybe someone smarter than me can chime in.

        • Uairhahs@lemmy.world
          link
          fedilink
          English
          arrow-up
          3
          ·
          9 hours ago

          Yes exactly the point being made in the link. AI picks up unintended behaviours already, it’s not too hard to imagine that the general direction is towards model collapse.

  • meme_historian@lemmy.dbzer0.com
    link
    fedilink
    English
    arrow-up
    244
    arrow-down
    1
    ·
    1 day ago

    Bruh, this is literally like scientists scavenging metal from shipwrecks that predate the era of nuclear bomb testing, cause they need steel that isn’t contaminated by nuclear fallout for certain measuring equipment 😵‍💫😵‍💫

    Great analogy for the shit we’re currently saturating our digital world with

    • isleepinahammock@lemmy.blahaj.zone
      link
      fedilink
      English
      arrow-up
      18
      arrow-down
      1
      ·
      12 hours ago

      Bruh, this is literally like scientists scavenging metal from shipwrecks that predate the era of nuclear bomb testing, cause they need steel that isn’t contaminated by nuclear fallout for certain measuring equipment

      Except in this case, these scientists are working at a nuclear weapons lab. They need metal that hasn’t been contaminated by nuclear bombs so that they can invent better nuclear bombs.

    • Courtney (she/her/they) @lemmy.blahaj.zone
      link
      fedilink
      English
      arrow-up
      28
      ·
      19 hours ago

      A guy I used to know who would go to estate sales looking for blacksmithing equipment once told me the market for old metalworking stuff is disappearing because too much stuff is being bought by companies én masse for this reason.

      I don’t know if I believe that companies are buying small lot auctions and estate sales for old anvils, but he was sure convinced.

      I like to joke I’m sitting on a pile of cash because my anvil isn’t radioactive. (it’s not a very good joke, but it carries some weight!)

      • BlaestEgnen@feddit.dk
        link
        fedilink
        English
        arrow-up
        5
        ·
        10 hours ago

        If a big company needed a lot of something specific, they’d probably pay a pretty penny and let people going to small lot auctions accumulate a sizeable amount and then buy it from them.

        It’s not impossible to buy stuff from small lots at scale, as long as you’re willing to pay a premium

        • Courtney (she/her/they) @lemmy.blahaj.zone
          link
          fedilink
          English
          arrow-up
          3
          ·
          3 hours ago

          My brother in law actually tried to hook me up at the car place he worked at, and my job would have been to go to auctions and estate sales looking for nice cars for the dealership to buy. It would essentially have been a couple things they want me to go to per week, but then I’d have to be looking myself for things like Craigslist, marketplace, and paper listings.

          Didn’t end up going for the job, due to the owners being shit head MAGAts, but it would have been interesting for a few months.

          That’s what I imagine any place that NEEDS pre-40s metal would do. Have a person who just goes to places that have blacksmithing tools advertised and buy them. Probably big auctions though, I imagine smallest ate sales wouldn’t get much attention.

      • WesternInfidels@feddit.online
        link
        fedilink
        English
        arrow-up
        25
        ·
        edit-2
        18 hours ago

        According to Wikipedia, while a market for old, low-background steel was real phenomenon for a few decades, that’s pretty much over.

        World anthropogenic background radiation … began to fall in 1963 … and by 2008 it had decreased to only 0.005 mSv/yr above natural levels.[9] This has made special low-background steel no longer necessary for most radiation-sensitive uses, as new steel now has a low enough radioactive signature.

        So you may consider your friend’s estate-sale theory in light of that time-line.

        but it carries some weight

        Did you know about !dadjokes@lemmy.world?

  • minus@lemmy.zip
    link
    fedilink
    English
    arrow-up
    108
    arrow-down
    4
    ·
    1 day ago

    The worst part is that they are destroying the books afterwards.

    • crusa187@lemmy.ml
      link
      fedilink
      English
      arrow-up
      58
      arrow-down
      1
      ·
      22 hours ago

      We also have no guarantee these books will be available at a later date in these digitized formats, nor that they won’t be altered to fit certain narratives which is of course trivial to do with digital copies. At a minimum these need public verifiable checksums to be relied upon.

      I don’t like this and don’t trust it one bit.

      These AI companies already showed their hand in trying to corner the PC component market with an end goal of users no longer being able to own their own PCs. Now they’re potentially trying to do this with all written recorded information.

      How long until Altman’s Firemen show up at your door to burn your books?

      • morto@piefed.social
        link
        fedilink
        English
        arrow-up
        22
        ·
        20 hours ago

        They even hve an economical incentive to destroy the books, because by doing so, they will get the data and prevent the competition from getting it for them

        • BlaestEgnen@feddit.dk
          link
          fedilink
          English
          arrow-up
          2
          ·
          10 hours ago

          The incentive is literally that it’s cheaper to scan, that’s all.

          But given they’re only interested in books with ISBN’s, I recon they’re already scanned and available in digital forms. If they’re newer than the 80s, they’re most likely in a document file in a server related to the publisher.

          Honestly surprised they’d rather source physical books, than just request various publishers for a price for their combined works - Only reason I can think of, is the physical scan makes other scans like a mobile photo easier to recognise and respond correctly to

    • AbouBenAdhem@lemmy.world
      link
      fedilink
      English
      arrow-up
      12
      ·
      20 hours ago

      Reminds me of Judge Holden in Blood Meridian, who sketched ancient artifacts in his personal notebook and then destroyed the originals.

      • Omgpwnies@lemmy.world
        link
        fedilink
        English
        arrow-up
        11
        arrow-down
        1
        ·
        19 hours ago

        More that they are destroying the books in the process of digitizing them, they basically cut the pages out of the book and feed them into a scanner with a document feeder attached.

        • T156@lemmy.world
          link
          fedilink
          English
          arrow-up
          16
          ·
          18 hours ago

          It’s worth pointing out that this is standard practice when digitising books in most places, if they’re not irreplaceable. Mass-produced books would fall under that category.

          Turning the page and photographing what’s there is really only done for some books, which you can’t afford to destroy.

          • Omgpwnies@lemmy.world
            link
            fedilink
            English
            arrow-up
            5
            arrow-down
            2
            ·
            17 hours ago

            From what I’ve read previously, they’re destroying all the books, regardless if they’re historic and one-of-a-kind or a mass produced paperback.

            • T156@lemmy.world
              link
              fedilink
              English
              arrow-up
              6
              ·
              13 hours ago

              At least in this case, yes. They’d just not get them, since they’d want something they can already easily process, and the source for the quote is a bookseller, who isn’t likely to sell something like that to begin with.

              An old book that might disintegrate and damage the scanner and hold things up would be undesirable, compared to either a new one, or a reprint of an old one.

              • cavitationfetishist2@quokk.au
                link
                fedilink
                English
                arrow-up
                1
                arrow-down
                2
                ·
                edit-2
                13 hours ago

                That’s not how most of that works.

                Things would be sold in bulk. Mistakes happen. Literally the only part of what you said that makes any sense is that fucked up pages could clog the pipe.

    • Riskable@programming.dev
      link
      fedilink
      English
      arrow-up
      16
      arrow-down
      33
      ·
      24 hours ago

      Is it really destroying though? They’re digitizing them, and publishers still have the digital copies ready to print more at any time. So it’s not like they’re destroying the texts, they’re just shifting them.

      Nobody complained when Google did this over a decade ago 🤷

      When you say they’re “destroying the books” you make it sound like they’re erasing one of the last known copy of some important work when in reality, most of these books were purchased in bulk from bookstores and libraries that were planning on discarding them anyway.

      Almost all these books were either headed to the dump or the recycling center. They’re just being digitized on the way.

      • Hawke@lemmy.world
        link
        fedilink
        English
        arrow-up
        57
        arrow-down
        1
        ·
        24 hours ago

        Digitized for private consumption.

        Nobody cares when Google did this or when archive.org does this, because they’re sharing the results with the world. (Idiotic shortsighted lawsuits from the authors guild notwithstanding)

        • antonim@lemmy.world
          link
          fedilink
          English
          arrow-up
          5
          ·
          edit-2
          16 hours ago

          Google (and/or the libraries that collaborate with Google Books) and Archive.org usually scan non-destructively. Some of Google’s book scans still show the fingers holding the corners of the pages, some of them were taken mistakenly mid-flip, etc. So they clearly used whole, normally bound books.

        • SkaveRat@discuss.tchncs.de
          link
          fedilink
          English
          arrow-up
          12
          ·
          23 hours ago

          in this case they often are actually destroying them. they take them apart, beause it’s easier to scan than to use a proper book scanner

          • Hawke@lemmy.world
            link
            fedilink
            English
            arrow-up
            5
            arrow-down
            1
            ·
            22 hours ago

            Yeah that’s true of Google as well I believe. I think Archive is more careful, but I’m sure some books get damaged in the process there as well.

            • Zarobi@aussie.zone
              link
              fedilink
              English
              arrow-up
              4
              ·
              12 hours ago

              Having done a lot of scanning in the past, no matter how gentle you are, the book will be damaged; at least a little bit. Even if you do it manually, by hand, really slowly.

              It’s the unfortunate reality, because books are not designed to be held open and pressed flat onto an unyielding surface. If you don’t press flat, you’ll get curved pages and bad quality (dark) images, which is sometimes ok and fixable by software, but often not. You can get really fancy expensive machines, but they still don’t fully solve the problem.

              I did scanlation editing stuff for a while, and half my job was making the raw page scans look presentable, because the scanners didn’t want to destroy their manga (understandable).

              https://en.wikipedia.org/wiki/Book_scanning#Methods

              • smh@slrpnk.net
                link
                fedilink
                English
                arrow-up
                3
                ·
                12 hours ago

                This matches my experience. My college archive digitizes a few yearbooks a year (for whatever classes are having as big reunion) and it’s a pain. For modern yearbooks (70s and later) we have a dozen or more copies, so we take one apart and feed it through the feed scanner, which works well. We then store the loose pages in a folder, in case we need to rescan anything.

                Older yearbooks (where we don’t have so many copies) we use the flatbed scanner and the pages come out warped, so we hand fix each page in Photoshop.

                • Zarobi@aussie.zone
                  link
                  fedilink
                  English
                  arrow-up
                  1
                  ·
                  11 hours ago

                  Yepp… memories of using the light level curve tool and white point and black point and the perspective tool. Then going in with the white brush and deleting any specks of dust that remained, and comparing to the original to make sure no lines got deleted

      • UnderpantsWeevil@lemmy.world
        link
        fedilink
        English
        arrow-up
        15
        ·
        edit-2
        20 hours ago

        Almost all these books were either headed to the dump or the recycling center.

        My wife works at a law firm that contracts with “Books By The Yard”, which provides books purely for office aesthetics. For a few hundred bucks you can plaster a bookshelf full of material nobody will ever read, because they’re such a commodity.

        It’s so crazy to see people ingesting this news from the conspiracy-brain perspective of “The AI companies are stealing all the knowledge!” without recognizing the more pressing reality of decade upon decade of publisher overproduction, resulting in a total devaluing of physical media for its academic importance.

        Imagine going into hysterics because a warehouse full of shitty airport books went up in smoke, like it’s the Library of Alexandra that just burned down.

        • isleepinahammock@lemmy.blahaj.zone
          link
          fedilink
          English
          arrow-up
          4
          ·
          12 hours ago

          Are you sure they’re digitizing Book By the Yard quality material?

          Consider this. Purchasing, shipping, and physically scanning a book is the single most expensive way for an AI company to acquire the text of a work. It’s been widely documented in court cases that these companies engage in mass IP theft. They’re literally running massive torrenting farms, grabbing copies of every film, song, book, etc. that they can get their hands on.

          What kinds of texts are most likely to be found pirated on the internet? It’s the mass market stuff. The common stuff was digitized long ago. They can just download that. They can pirate it. They can buy the ebook. There’s no need for OpenAI to purchase and destructively scan the works of Steven King. No shade on the man, but his works aren’t exactly hard to find. I’m sure they can just find a torrent.

          There’s little value in scanning the mass market books that are produced in enormous quantities. They probably don’t buy “Book by the Yard” books, because they already have digital copies of those.

          But the rare, long out-of-print stuff? The stuff that you actually cannot find a legal or illegal copy of online? That’s only stuff worth paying money to buy, ship, and scan.

          Doing anything physical, especially at scale, is slow and expensive. I would guess a good portion of these books, perhaps an outright majority, have simply never been digitized.

          • UnderpantsWeevil@lemmy.world
            link
            fedilink
            English
            arrow-up
            1
            ·
            2 hours ago

            re you sure they’re digitizing Book By the Yard quality material?

            They were more than happy to gobble up Reddit posts and Twitter farts and the worst of 4chan. You don’t need to ingest Ulysses exclusively to train a system.

            Doing anything physical, especially at scale, is slow and expensive.

            Which is why you get people paid pennies an hour to break the physical copies down and feed them into scanners.

            would guess a good portion of these books, perhaps an outright majority, have simply never been digitized.

            Possibly not. But that’s a sign it isn’t great literature or coveted texts.

            • isleepinahammock@lemmy.blahaj.zone
              link
              fedilink
              English
              arrow-up
              1
              ·
              35 minutes ago

              You’re confusing uniqueness with quantity. 4chan posts are not Ulysses, but they are still useful data.

              My point was not that they would refuse to scan Book By the Yard quality works because they would consider them unsuitable. My point was that they don’t need to scan mass market Book by the Yard stuff, as they already have digital copies of it.

              And scanning books is not as cheap as you think it is. Even using labor in low income countries, the cost of scanning is still vastly greater than just downloading a file.

      • minus@lemmy.zip
        link
        fedilink
        English
        arrow-up
        23
        arrow-down
        1
        ·
        edit-2
        22 hours ago

        When they destroy the physical copy they remove them from the antique stores market which often rely on circulation.

        Edit: Also physical copies don’t require electricity, a device and internet access plus they are something you can own.

        • Riskable@programming.dev
          link
          fedilink
          English
          arrow-up
          6
          ·
          edit-2
          18 hours ago

          If you think any more than 0.1% of these physical books would ever have ended up in antique bookstores, you’re dreaming.

          Think about how many books out there are things like Donald Trump’s biography, or pointless drivel from non-experts, self-help books that tell people to down “essential oils”, old editions of programming books, or just plain shitty fiction that never sold much in the first place.

          It’s ok to throw trash away! Really!

          • isleepinahammock@lemmy.blahaj.zone
            link
            fedilink
            English
            arrow-up
            2
            ·
            12 hours ago

            Why would they scan that stuff though? That kind of mass market (human made) slop is easily available in digital form. We know the AI companies have engaged in mass digital piracy, including running massive torrenting operations. So if a digital copy exists, they probably already have it. And even if they have to buy it, purchasing an ebook is a lot cheaper than buying a physical one, shipping it, paying someone to scan it, etc.

            I would think the old, the out-of-print, the rare, and never-before digitized are the only things worth buying and physically scanning in 2026. Everything else has already been scanned or was born digital-native.

        • XLE@piefed.social
          link
          fedilink
          English
          arrow-up
          3
          arrow-down
          1
          ·
          22 hours ago

          At least digital storage prices have only been going down in recent months, right? /s

        • Snot Flickerman@lemmy.blahaj.zone
          link
          fedilink
          English
          arrow-up
          24
          arrow-down
          1
          ·
          23 hours ago

          The knowledge is retained and owned by a private corporation who now does not have to share what may have been a still under copyright, but now the corporation owns it? Make it make sense.

          • Riskable@programming.dev
            link
            fedilink
            English
            arrow-up
            2
            ·
            19 hours ago

            Yes. That does make sense.

            If you bought a book, scanned it—destroying it in the process—then read it on your computer, that would be completely acceptable.

            Why is it wrong when a corporation does the same thing?

            They’re not claiming ownership of the copyrights, just ownership of a copy. Which is how copyright works.

            • BlaestEgnen@feddit.dk
              link
              fedilink
              English
              arrow-up
              1
              ·
              10 hours ago

              Now we just need to proof they’ve copied the files onto a second hard drive, so they own two copies

              If we’re gonna keep it as a copyright issue, I think people are mostly mad they’re not able to reprint the books themselves if they wanted to. We’re not speaking from a copyright position

  • kescusay@lemmy.world
    link
    fedilink
    English
    arrow-up
    47
    ·
    1 day ago

    I’m waiting for the day when desperate LLM companies start paying people to post real human content, only for those people to just ask ChatGPT to do it.

    • Meron35@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      ·
      3 hours ago

      Already happening for years.

      Amazon mechanical turk, Outlier.ai, scale ai, are all gig platforms that pay people to produce real human content for LLM training.

      At first they were paying people to actually produce content, like worked solutions to maths problems or translations, then they pivoted more to rating LLM output.

      And as with every enshittification cycle, wages and work rapidly dried up, so people responded by asking LLMs for the answer just to meet deadlines and get by.

    • Tollana1234567@lemmy.today
      link
      fedilink
      English
      arrow-up
      1
      ·
      9 hours ago

      more than likely they are scraping other AI bot posts, there isnt that many users on most of the social media sites. its 50/50 AI/BOTS for most of them.

    • isleepinahammock@lemmy.blahaj.zone
      link
      fedilink
      English
      arrow-up
      1
      ·
      11 hours ago

      I envision the opposite:

      In the future, I imagine authors will go to great lengths to prove their works to be authentically human. It may become standard practice for authors to live-stream their writing. Literally a years-long stream of them sitting at their desk. Comments turned off.

      Perhaps libraries could even offer this as a service. Offer a space where authors can come write in public and have it simultaneously live-streamed. Maybe it becomes the norm for authors to mostly write publicly on a platform like Twitch, but to occasionally write in public spaces where they can be physically observed.

      This will not be a legal requirement to publish a book. But no respectable author would do otherwise.

      • jaygray91@piefed.zip
        link
        fedilink
        English
        arrow-up
        2
        ·
        8 hours ago

        I imagine authors will go to great lengths to prove their works to be authentically human

        artists (the drawing kind) have been posting timelapse of them drawing art for quite some time. It has since increased when AI sloperators flooded the field. But now I’ve heard some aren’t even going to post them anymore, since the sloperators are training on those to make fake timelapse.

    • veryblandusername@fedinsfw.app
      link
      fedilink
      English
      arrow-up
      8
      ·
      20 hours ago

      This is already happening.

      They are paying people bottom market rates to generate unique written content and people are just having AI write it for them.

    • pdxfed@lemmy.world
      link
      fedilink
      English
      arrow-up
      10
      ·
      1 day ago

      “now you only posted 4 times today Jenny, do you just want to do the minimum?”

    • Riskable@programming.dev
      link
      fedilink
      English
      arrow-up
      3
      arrow-down
      4
      ·
      24 hours ago

      Big AI mostly switched to synthetic training data anyway. The books they’re digitizing are being used to gather knowledge, not writing styles or logic (mostly).

      As in, when you ask ChatGPT how long some book is, it can just go check (if it’s in the database). It’s also useful if you ask about that book or about knowledge contained in that book. It’ll even reference books now (if you demand that in your prompt).

      It’s not the same as earlier LLM tech which relied on scanned text to figure out how to respond to any given prompt (from a language standpoint). The “language” part of LLMs is a solved problem now (thanks to the synthetic training). At least for English 🤷

      • AbouBenAdhem@lemmy.world
        link
        fedilink
        English
        arrow-up
        6
        ·
        22 hours ago

        The article quotes a post from ISBNdb saying the issue is model collapse from training on synthetic data.

        • Riskable@programming.dev
          link
          fedilink
          English
          arrow-up
          2
          arrow-down
          1
          ·
          19 hours ago

          That’s like saying, “they had some failure modes from the synthetic data, so they should just obviously stop trying forever.

          They’ll just fix the edge cases and move on. Like any programming task.

          • kescusay@lemmy.world
            link
            fedilink
            English
            arrow-up
            2
            arrow-down
            1
            ·
            16 hours ago

            The problem is that synthetic data is not fit for that purpose. The more of it you use, the worse at dealing with the cases LLMs get.

            Think of it like this… You feed a language model a bunch of genuine human-written content. Great. Now it can produce the most likely text in a lot of cases. Word combinations that rarely appear in written language rarely get generated, so most of its synthetic data lacks those rare - but still valid - combinations.

            Train it on this synthetic data, and now more outliers and rare combinations get filed off. Rinse and repeat.

            • isleepinahammock@lemmy.blahaj.zone
              link
              fedilink
              English
              arrow-up
              1
              ·
              10 hours ago

              I think what they’re saying is that the AI companies don’t need to rinse and repeat. They’ve already created solid programs for how to make LLMS speak English and demonstrate basic reasoning. You don’t need to retrain that part every time. Once you’ve trained “English.exe,” you can just copy it endlessly.

              Maybe with enough time, the hard-coded English of the LLMs could become increasingly anachronistic and sound old-fashioned and formal to most ears. But for that kind of for slow maintenance you could just pay people to write examples of modern language and train it on that.

              • kescusay@lemmy.world
                link
                fedilink
                English
                arrow-up
                1
                ·
                3 hours ago

                Then that’s a misunderstanding on their part. What I’m getting at is that if all of your new training data doesn’t reinforce uncommon - but factually and grammatically correct - outlier word relationships, then those outliers fade away.

                There’s also an amplification issue. OpenAI has had to add a ton of instructions to their harnesses not to mention goblins because the model trained on a bunch of synthetic data when one of ChatGPT’s offered personalities would go “goblin mode.” So the more it mentioned goblins, the more that data was accidentally fed back into it, and all of a sudden no matter which personality you assigned, ChatGPT would go off about goblins.

    • luciferofastora@feddit.org
      link
      fedilink
      English
      arrow-up
      2
      ·
      11 hours ago

      Don’t you know that the solution to any and all problems with AI just need more training data to fix? If we can teach it to be just a little more human, it’ll surely stop going off the rails. We just need to feed the machine more data More Data MORE DATA MORE DATA

  • SparroHawc@piefed.world
    link
    fedilink
    English
    arrow-up
    18
    ·
    24 hours ago

    Nice that they’re actually buying them this time instead of just downloading torrents of book collections again.

        • isleepinahammock@lemmy.blahaj.zone
          link
          fedilink
          English
          arrow-up
          2
          ·
          10 hours ago

          Not the rare, out of print, and long forgotten. And that’s what’s so scary about this. They’re probably not paying to scan books that would otherwise end up at a landfill.

        • zerofk@lemmy.zip
          link
          fedilink
          English
          arrow-up
          1
          ·
          9 hours ago

          Far from it. I think you vastly underestimate how many books have been published, and in what small quantities.

          Estimates are around 15% of books is digitised, and less than a third of that is searchable. I’m guessing a small fraction of that is torrentable (yes that’s a word I just made up).

    • Axolotl@feddit.it
      link
      fedilink
      English
      arrow-up
      7
      arrow-down
      1
      ·
      20 hours ago

      They first crashed RAM prices, the SSDs, and now books too?! Oh fuck 😭

  • FlashMobOfOne@lemmy.world
    link
    fedilink
    English
    arrow-up
    3
    ·
    18 hours ago

    Yeah, pretty awful on so many different levels.

    Brilliant decision to train their models on the worst content to begin with.