You can run it on your desktop if you have a shit ton of vram. There are more compact versions, but the full model is 55.6 GB. You could run it on an Nvidia H100 80GB, which runs for $40k. Or more likely you’d run it on a Mac Studio M3 Ultra with 96GB of unified memory, which still runs for $4k.
im running it on my desktop with about 30 t/s on a computer from 2 years ago with a tpu bolted on. im really good at researching things, its my job. i do science™️. its a little worse at researching things online than i am and rarely hallucinates. its especially good with unlimited search toolcalls. i was able to debug a problem in my code that would’ve taken me all day to figure out but after quite a lot of wiresharking and pushing logs to qwen i was able to resolve the problem in 45 minutes. its also often able to find unpaywalled scientific articles that are hidden in every major search engine which is very useful when anna’s archive can’t pull through.
these things are getting fairly impressive to me and i am generally very skeptical of them and skim everything they do. qwen is doing some tremendous work especially in democratizing them as general use tools. 30 t/s isn’t great, I have to wait quite a bit for it to think, but its better than feeding a ton of my data to some lab somewhere.
This is really neat to hear. Can I ask which version of the model you’re running? And which TPU you’re using? If it’s usable for work I’d be happy to poke my bosses about letting me use a local model.
i’ve been using qwen 3.6 35b a3b, 64gb ram, a used 4080, llama.cpp with some tweaks for more tool calling, and some amd cpu that has a npu in it from a couple years back i cant be assed to boot it up right now to check lol
i got the ram and gpu before everything went completely to shit, whole setup was like 2k usd or something but was kinda important to my job so i got it with a partial rebate from my employer. i havent really tried to optimize it much. qwen 35b a3b has slightly worse performance than 27b but runs a lot faster.
The cheapest way to run larger models locally is the Vega architecture Radeon Instinct MI60. 32gb of HBM for 3-500 usd used, add 30-40 per unit for active cooling.
You can run it on your desktop if you have a shit ton of vram. There are more compact versions, but the full model is 55.6 GB. You could run it on an Nvidia H100 80GB, which runs for $40k. Or more likely you’d run it on a Mac Studio M3 Ultra with 96GB of unified memory, which still runs for $4k.
im running it on my desktop with about 30 t/s on a computer from 2 years ago with a tpu bolted on. im really good at researching things, its my job. i do science™️. its a little worse at researching things online than i am and rarely hallucinates. its especially good with unlimited search toolcalls. i was able to debug a problem in my code that would’ve taken me all day to figure out but after quite a lot of wiresharking and pushing logs to qwen i was able to resolve the problem in 45 minutes. its also often able to find unpaywalled scientific articles that are hidden in every major search engine which is very useful when anna’s archive can’t pull through.
these things are getting fairly impressive to me and i am generally very skeptical of them and skim everything they do. qwen is doing some tremendous work especially in democratizing them as general use tools. 30 t/s isn’t great, I have to wait quite a bit for it to think, but its better than feeding a ton of my data to some lab somewhere.
This is really neat to hear. Can I ask which version of the model you’re running? And which TPU you’re using? If it’s usable for work I’d be happy to poke my bosses about letting me use a local model.
i’ve been using qwen 3.6 35b a3b, 64gb ram, a used 4080, llama.cpp with some tweaks for more tool calling, and some amd cpu that has a npu in it from a couple years back i cant be assed to boot it up right now to check lol
i got the ram and gpu before everything went completely to shit, whole setup was like 2k usd or something but was kinda important to my job so i got it with a partial rebate from my employer. i havent really tried to optimize it much. qwen 35b a3b has slightly worse performance than 27b but runs a lot faster.
Running fp16, in current year?
https://huggingface.co/unsloth/Qwen3.6-27B-GGUF
even 5ks is 19gb. Plus context but that still fully fits on a 3090. But even splitting it across gpu + cpu isn’t as bad as it could be.
The cheapest way to run larger models locally is the Vega architecture Radeon Instinct MI60. 32gb of HBM for 3-500 usd used, add 30-40 per unit for active cooling.