• brucethemoose@lemmy.world
    link
    fedilink
    English
    arrow-up
    6
    arrow-down
    1
    ·
    edit-2
    8 hours ago

    Well unfortunately for those investors, they aren’t giving their money to actual AGI research, but to scammers trying to sell infinite scaling of transformers… which has nothing to do with AGI.

    • Nouvellalia@lemmy.world
      link
      fedilink
      English
      arrow-up
      3
      arrow-down
      1
      ·
      edit-2
      7 hours ago

      Transformers are only a part of AGI, just like “saying the most logical next word” is only a part of the human mind. Transformers, which need not be used exclusively on language, are a great way to dial in successful actions from the infinity of possibility. They are not a great way of vetting those actions before you implement them. They are a necessary part of AGI, just like the human use of language was a necessary part of us getting to the era of space ships.

      I don’t think investors are being sold the idea that infinite scaling transformers are the end of the road. From what I’ve seen, anthropic and openai have both been desperately trying to add systems to their transformers to make them more capable, while also continuing to scale transformers to find out where that particular technology plateaus. This has been their path since GPT4 was released.

      I think your assessment is an extreme oversimplification which naturally looks like it will fail because you left out 90% of what is going on. Are you only following the media meant for the general public, and the PR statements from AI companies? Like most scientific or businesses endeavors, that’s been dumbed down to the point of being useless, just so the average person can understand what’s happening, or is simply advertising.

      • brucethemoose@lemmy.world
        link
        fedilink
        English
        arrow-up
        2
        ·
        edit-2
        5 hours ago

        Okay.

        More specifically, autoregressive transformers LLMs are not a path to a component of AGI.

        The architecture is absolutely terrible for such a thing, for so many reasons. I don’t know how anyone who’s played with them can say otherwise and believe it; it’s like saying blimps are a viable path to the moon. It has its niches, but AGI is not one of them.

        Calling them a stepping stone is a… stretch.

        Maybe world models that “train as they go” and have long moved on from transformers are a bit closer, like a few researchers are playing up, but again… that has almost nothing to do with transformers LLMs. Its why researchers distanced themselves from that.

        I think your assessment is an extreme oversimplification which naturally looks like it will fail because you left out 90% of what is going on. Are you only following the media meant for the general public, and the PR statements from AI companies? Like most scientific or businesses endeavors, that’s been dumbed down to the point of being useless, just so the average person can understand what’s happening, or is simply advertising.

        I dunno why everyone always jumps to accusations like this.

        I’ve been hacking/toying with LLMs on my desktop since 2021, and with GANs and other models before that. I’ve done professional work with text models. I keep up with papers, on-and-off, and upload experiments. I’m not a researcher or anything; I’m just a hobbyist.

        But I’m certainly not following any AI YouTubers or anything like that. And I wouldn’t trust Sam Altman if he said the sky is blue.

        • Nouvellalia@lemmy.world
          link
          fedilink
          English
          arrow-up
          1
          ·
          edit-2
          3 hours ago

          I dunno why everyone always jumps to accusations like this.

          Well, I can’t speak to everyone, but I jumped to that conclusion because you proposed the idea that the current top of the line models are simply big fat transformers and nothing else.

          I’m really glad to hear that you’ve had your fingers in the pie, and have a good grasp of what an LLM is and what it does on a mechanical level. It makes talking about them much easier. I’m super cool with continuing a discussion if you are interested in why I think they are an integral part of eventual AGI.

          To clarify, I do not think that their current form is a 1:1. I disagree with your blimp analogy though. I think a better comparison would be an internal combustion engine. It’s certainly not a turbofan engine or a scramjet, but the basic concept is there.

          Turning language into a fuzzy world model that can be interfaced with simply, and performing math on that model, is just as important to AGI as a fully fleshed out physics simulation is for AGI. Humans have both. Why would an AGI not need them?

          What an LLM does to it’s array is not “intelligence”. It’s simple math. However, it is performed on a thing so complex that it cannot be derived from any math we have now, a manifold created through eons of interface between all humans and the universe.

          The “magic” isn’t in the simple math, it’s in the thing we all made together for millions of years. An interface with this, even one as “simple” as an autoregressive transformer, is an indespensible part of AGI, both because it provides a fuzzy way to understand the universe, and because it provides an interface with humanity, and humanity processes the universe in a way that is so complex, computers will not be able to emulate it for generations.

          Edit: also, fuck altman and musk and zuck and even amodei. I wouldn’t trust anything they say either.