• kescusay@lemmy.world
    link
    fedilink
    English
    arrow-up
    47
    ·
    1 day ago

    I’m waiting for the day when desperate LLM companies start paying people to post real human content, only for those people to just ask ChatGPT to do it.

    • Meron35@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      ·
      3 hours ago

      Already happening for years.

      Amazon mechanical turk, Outlier.ai, scale ai, are all gig platforms that pay people to produce real human content for LLM training.

      At first they were paying people to actually produce content, like worked solutions to maths problems or translations, then they pivoted more to rating LLM output.

      And as with every enshittification cycle, wages and work rapidly dried up, so people responded by asking LLMs for the answer just to meet deadlines and get by.

    • Tollana1234567@lemmy.today
      link
      fedilink
      English
      arrow-up
      1
      ·
      9 hours ago

      more than likely they are scraping other AI bot posts, there isnt that many users on most of the social media sites. its 50/50 AI/BOTS for most of them.

    • isleepinahammock@lemmy.blahaj.zone
      link
      fedilink
      English
      arrow-up
      1
      ·
      11 hours ago

      I envision the opposite:

      In the future, I imagine authors will go to great lengths to prove their works to be authentically human. It may become standard practice for authors to live-stream their writing. Literally a years-long stream of them sitting at their desk. Comments turned off.

      Perhaps libraries could even offer this as a service. Offer a space where authors can come write in public and have it simultaneously live-streamed. Maybe it becomes the norm for authors to mostly write publicly on a platform like Twitch, but to occasionally write in public spaces where they can be physically observed.

      This will not be a legal requirement to publish a book. But no respectable author would do otherwise.

      • jaygray91@piefed.zip
        link
        fedilink
        English
        arrow-up
        2
        ·
        8 hours ago

        I imagine authors will go to great lengths to prove their works to be authentically human

        artists (the drawing kind) have been posting timelapse of them drawing art for quite some time. It has since increased when AI sloperators flooded the field. But now I’ve heard some aren’t even going to post them anymore, since the sloperators are training on those to make fake timelapse.

    • veryblandusername@fedinsfw.app
      link
      fedilink
      English
      arrow-up
      8
      ·
      20 hours ago

      This is already happening.

      They are paying people bottom market rates to generate unique written content and people are just having AI write it for them.

    • pdxfed@lemmy.world
      link
      fedilink
      English
      arrow-up
      10
      ·
      1 day ago

      “now you only posted 4 times today Jenny, do you just want to do the minimum?”

    • Riskable@programming.dev
      link
      fedilink
      English
      arrow-up
      3
      arrow-down
      4
      ·
      24 hours ago

      Big AI mostly switched to synthetic training data anyway. The books they’re digitizing are being used to gather knowledge, not writing styles or logic (mostly).

      As in, when you ask ChatGPT how long some book is, it can just go check (if it’s in the database). It’s also useful if you ask about that book or about knowledge contained in that book. It’ll even reference books now (if you demand that in your prompt).

      It’s not the same as earlier LLM tech which relied on scanned text to figure out how to respond to any given prompt (from a language standpoint). The “language” part of LLMs is a solved problem now (thanks to the synthetic training). At least for English 🤷

      • AbouBenAdhem@lemmy.world
        link
        fedilink
        English
        arrow-up
        6
        ·
        22 hours ago

        The article quotes a post from ISBNdb saying the issue is model collapse from training on synthetic data.

        • Riskable@programming.dev
          link
          fedilink
          English
          arrow-up
          2
          arrow-down
          1
          ·
          19 hours ago

          That’s like saying, “they had some failure modes from the synthetic data, so they should just obviously stop trying forever.

          They’ll just fix the edge cases and move on. Like any programming task.

          • kescusay@lemmy.world
            link
            fedilink
            English
            arrow-up
            2
            arrow-down
            1
            ·
            16 hours ago

            The problem is that synthetic data is not fit for that purpose. The more of it you use, the worse at dealing with the cases LLMs get.

            Think of it like this… You feed a language model a bunch of genuine human-written content. Great. Now it can produce the most likely text in a lot of cases. Word combinations that rarely appear in written language rarely get generated, so most of its synthetic data lacks those rare - but still valid - combinations.

            Train it on this synthetic data, and now more outliers and rare combinations get filed off. Rinse and repeat.

            • isleepinahammock@lemmy.blahaj.zone
              link
              fedilink
              English
              arrow-up
              1
              ·
              10 hours ago

              I think what they’re saying is that the AI companies don’t need to rinse and repeat. They’ve already created solid programs for how to make LLMS speak English and demonstrate basic reasoning. You don’t need to retrain that part every time. Once you’ve trained “English.exe,” you can just copy it endlessly.

              Maybe with enough time, the hard-coded English of the LLMs could become increasingly anachronistic and sound old-fashioned and formal to most ears. But for that kind of for slow maintenance you could just pay people to write examples of modern language and train it on that.

              • kescusay@lemmy.world
                link
                fedilink
                English
                arrow-up
                1
                ·
                3 hours ago

                Then that’s a misunderstanding on their part. What I’m getting at is that if all of your new training data doesn’t reinforce uncommon - but factually and grammatically correct - outlier word relationships, then those outliers fade away.

                There’s also an amplification issue. OpenAI has had to add a ton of instructions to their harnesses not to mention goblins because the model trained on a bunch of synthetic data when one of ChatGPT’s offered personalities would go “goblin mode.” So the more it mentioned goblins, the more that data was accidentally fed back into it, and all of a sudden no matter which personality you assigned, ChatGPT would go off about goblins.