• RobotToaster@mander.xyz
    link
    fedilink
    English
    arrow-up
    10
    arrow-down
    1
    ·
    19 hours ago

    I can’t afford a $10k computer to run local models.

    Hopefully that is only temporary.

    I imagine few people could afford a Unix system in 1987 when Stallman first wrote GCC.

    • Snort_Owl [they/them]@hexbear.net
      link
      fedilink
      English
      arrow-up
      20
      ·
      edit-2
      18 hours ago

      It’s not magic these models are big and you can’t just squeeze them onto smaller devices. Quantisation isn’t magic either it’s literally reducing the accuracy of a larger model to fit on a smaller device reducing its performance. If the data doesn’t exist in the model weights then the model will perform worse and you can only compress data so much. The only meaningful tech shift in the space is unified memory but it’s not going to be enough for anything this is fundamentally going to need silly quantities of ram for any effective model to run.

      • Chana [none/use name]@hexbear.net
        link
        fedilink
        English
        arrow-up
        5
        ·
        11 hours ago

        The main limitation is just RAM on the GPUs. Unfortunately the bubble itself drove up prices on that exact thing dramatically but this not a particularly expensive thing normally.

        • ashinadash [she/her]@hexbear.net
          link
          fedilink
          English
          arrow-up
          7
          ·
          11 hours ago

          normally.

          GPU prices shot up in ~2018 and 2020 and [etc] and never came back down. Nowadays a 9060XT (the useful version, mind) costs on its own what a decent entry level PC used to. We are never going back to normal.


          Remember to enable JPEG-XL support in your browser, even on your phone!

    • insurgentrat [she/her, it/its]@hexbear.net
      link
      fedilink
      English
      arrow-up
      21
      ·
      19 hours ago

      We’re at a really different point technologically and given the whole thing of Google’s paper and the architecture is more data more layers more good there is a huge incentive to run on as much ram as possible which is expensive.

      We’re not going from hand soldered resistors to nanoscale transistors again. We’re looking at some serious physics walls. Even if photonic chips leave the lab.