• brucethemoose@lemmy.world
    link
    fedilink
    English
    arrow-up
    2
    ·
    edit-2
    vor 22 Stunden

    Even “big” open source models like DSV4 and Ling/Ring are very efficient. They’re big, but (seemingly) sparser than US models, so they’re cheap.

    They run surprisingly well with hybrid CPU+GPU inference on desktops. And thats not even getting into the efficient attention mechanisms.

    I can run DSV4 Flash, barely quantized, with ~1M context on my Ryzen desktop at ~11 tokens/s. If you told me that two years ago, I would not have believed you.