• MeetMeAtTheMovies [they/them]@hexbear.net
    link
    fedilink
    English
    arrow-up
    8
    ·
    4 days ago

    This is really neat to hear. Can I ask which version of the model you’re running? And which TPU you’re using? If it’s usable for work I’d be happy to poke my bosses about letting me use a local model.

    • lurkerlady [she/her]@hexbear.net
      link
      fedilink
      English
      arrow-up
      11
      ·
      edit-2
      4 days ago

      i’ve been using qwen 3.6 35b a3b, 64gb ram, a used 4080, llama.cpp with some tweaks for more tool calling, and some amd cpu that has a npu in it from a couple years back i cant be assed to boot it up right now to check lol

      i got the ram and gpu before everything went completely to shit, whole setup was like 2k usd or something but was kinda important to my job so i got it with a partial rebate from my employer. i havent really tried to optimize it much. qwen 35b a3b has slightly worse performance than 27b but runs a lot faster.