eicker@lemmy.world to Technology@lemmy.worldEnglish · 2 days agoOpenAI Hoarding Tens Of Thousands Of Apple Mac mini And Mac Studio Devices, As ASUS And MSI Burn Through Their Entire First Batch Of NVIDIA RTX Spark Chip And Beg For More.wccftech.comexternal-linkmessage-square80fedilinkarrow-up1332arrow-down11
arrow-up1331arrow-down1external-linkOpenAI Hoarding Tens Of Thousands Of Apple Mac mini And Mac Studio Devices, As ASUS And MSI Burn Through Their Entire First Batch Of NVIDIA RTX Spark Chip And Beg For More.wccftech.comeicker@lemmy.world to Technology@lemmy.worldEnglish · 2 days agomessage-square80fedilink
minus-squareLydia_K@lemmy.worldlinkfedilinkEnglisharrow-up8arrow-down1·2 days agohttps://github.com/AtomicBot-ai/atomic-llama-cpp-turboquant I’m running gwen 3.6 with 131k context window on a 3090, it’s fast enough and about as good as pay to play Claude at work.
minus-squareArchAengelus@lemmy.dbzer0.comlinkfedilinkEnglisharrow-up3·2 days agoUpgrade that to 3.8 as soon as your hardware allows (and your use case makes sense). 3.8 is quite a bit more rational.
minus-squareLydia_K@lemmy.worldlinkfedilinkEnglisharrow-up2arrow-down1·2 days agoI plan to once there is a version with turboquant and MTP as that huge context window is key.
minus-squareboonhet@sopuli.xyzlinkfedilinkEnglisharrow-up1·2 days agoA used 3090 is like 2-3k though, IF you can find one :| A month of claude is like 20 EUR. A month of opencode go is half that, but you get less usage.
https://github.com/AtomicBot-ai/atomic-llama-cpp-turboquant
I’m running gwen 3.6 with 131k context window on a 3090, it’s fast enough and about as good as pay to play Claude at work.
Upgrade that to 3.8 as soon as your hardware allows (and your use case makes sense). 3.8 is quite a bit more rational.
I plan to once there is a version with turboquant and MTP as that huge context window is key.
A used 3090 is like 2-3k though, IF you can find one :| A month of claude is like 20 EUR. A month of opencode go is half that, but you get less usage.