It’s not magic these models are big and you can’t just squeeze them onto smaller devices. Quantisation isn’t magic either it’s literally reducing the accuracy of a larger model to fit on a smaller device reducing its performance. If the data doesn’t exist in the model weights then the model will perform worse and you can only compress data so much. The only meaningful tech shift in the space is unified memory but it’s not going to be enough for anything this is fundamentally going to need silly quantities of ram for any effective model to run.
The main limitation is just RAM on the GPUs. Unfortunately the bubble itself drove up prices on that exact thing dramatically but this not a particularly expensive thing normally.
GPU prices shot up in ~2018 and 2020 and [etc] and never came back down. Nowadays a 9060XT (the useful version, mind) costs on its own what a decent entry level PC used to. We are never going back to normal.
ⓘ Remember to enable JPEG-XL support in your browser, even on your phone!
We’re at a really different point technologically and given the whole thing of Google’s paper and the architecture is more data more layers more good there is a huge incentive to run on as much ram as possible which is expensive.
We’re not going from hand soldered resistors to nanoscale transistors again. We’re looking at some serious physics walls. Even if photonic chips leave the lab.
Hopefully that is only temporary.
I imagine few people could afford a Unix system in 1987 when Stallman first wrote GCC.
It’s not magic these models are big and you can’t just squeeze them onto smaller devices. Quantisation isn’t magic either it’s literally reducing the accuracy of a larger model to fit on a smaller device reducing its performance. If the data doesn’t exist in the model weights then the model will perform worse and you can only compress data so much. The only meaningful tech shift in the space is unified memory but it’s not going to be enough for anything this is fundamentally going to need silly quantities of ram for any effective model to run.
The main limitation is just RAM on the GPUs. Unfortunately the bubble itself drove up prices on that exact thing dramatically but this not a particularly expensive thing normally.
GPU prices shot up in ~2018 and 2020 and [etc] and never came back down. Nowadays a 9060XT (the useful version, mind) costs on its own what a decent entry level PC used to. We are never going back to normal.
ⓘ Remember to enable JPEG-XL support in your browser, even on your phone!
Excuse my ma’am but let me tell you about a little thing called supply and demand. smirks insufferably just before getting punched
We’re at a really different point technologically and given the whole thing of Google’s paper and the architecture is more data more layers more good there is a huge incentive to run on as much ram as possible which is expensive.
We’re not going from hand soldered resistors to nanoscale transistors again. We’re looking at some serious physics walls. Even if photonic chips leave the lab.