• pongo1231 [none/use name]@hexbear.net
    link
    fedilink
    English
    arrow-up
    8
    arrow-down
    1
    ·
    edit-2
    13 hours ago

    I can’t afford a $10k computer to run local models.

    Tbh you don’t really need to either. With models like Qwen 3.6 35B A3B (which is quite close to the performance of frontier models from a year ago) enough of it fits onto a 8GB GPU - with the rest sitting in RAM or swapped out - to have it running at a speed that is very workable with. On my work laptop with a 4060 equivalent and llama.cpp + CUDA I get around 15-20 t/s for coding tasks. The equivalent Qwen 3.8 variant should be even better once that releases. I strongly believe small models like these running locally is going to be the way to go once the bubble bursts.